A developer on our team spent an afternoon getting a bug-report form to file tickets into a client's ticketing system. It worked. The agent filed three tickets, then queried the system and read every field back: title, description, steps to reproduce, and the reproducibility dropdown. Each field was correct. The agent reported success, and the report was true.
The feature was still broken.
The form created tickets of the wrong type. The workflow for that type marks an item resolved the moment it arrives, so every new bug report arrived already closed, where nobody filtering for open work would see it. The agent had noticed the symptom. It wrote that a newly filed bug in a resolved status would not appear in open-work filters. Then it blamed the client's configuration and suggested raising it with them.
The client's configuration was fine. The issue type was set in our code.
Nothing here was a coding problem
Every line the agent wrote was correct. Every check it ran passed, and the passes were real. If you had graded that session on what we have graded developers on for decades, whether the code works, it would have scored full marks.
What was missing was someone who doubted that "it works" was the whole answer.
That kind of doubt made a good engineer in 1998, and it matters just as much now that a machine does the typing. The list of things that make someone good at this job has barely changed. What has changed is that the skill we always taught first and hired for, writing code that works, is now the one an agent does well. Whole interview formats exist to test that one skill, and companies still use them.
Five skills did not get automated. "Critical thinking" is too broad to teach, but these five can be taught. The examples below are all ours: one junior developer's afternoon, and two months of transcripts from the senior engineer on the same codebase.
The five
Breaking a problem into parts that can each be wrong on their own. The goal is that when something is wrong, you can tell which part. In our transcripts, the senior version of this is an instruction sent before any work started: "Do phase 1 only: bump every place the version is pinned, and tell me what a v2 migration would cost here before doing it." Two pieces, one of them held back, with a question attached to the one held back. Nothing stops an agent from working this way. Someone has to ask it to.
Trying to prove your own answer wrong. The strongest example in two months of transcripts is eleven words, typed at three in the morning and never edited:
you can check if it does. i doubt it. but check.
That message holds two instructions: test this, and here is what I already believe. Ask an agent "how should we do this?" and it produces something reasonable, then agrees with itself. Give an agent a belief and tell it to test it, and it has something to push against. The output is completely different, because a person held an opinion first.
Not accepting the first plausible answer. This reflex is separate from the last one. It comes when the answer arrives, where the last one comes before. "what? that makes no sense, every sprint before this, we've been squash merging and all the details were in the PR description. why would we change that now? explain please." The person asking did not know the answer. They noticed that it contradicted something they knew, and said so before moving on.
Saying what you expect before the work starts. Our best example of this came from the junior developer. Before touching anything, she wrote: "all i need to verify is that on localhost i go to report a bug and submit and it gets submitted in jira." That is a clear acceptance test, written in advance, in plain language. She already had this skill.
Knowing where the checking stops. An agent checks the layer it can reach. In the ticket story, that layer was whether the interface accepted the data and whether the values arrived intact. Both were true, and both were the wrong question. Whether a bug report is the right kind of ticket to create at all is a decision about what the feature is for, and reading the response back will never reveal it. The limit moves with the task. An agent that can run the test suite will confirm that the tests pass. That says nothing about whether the tests check the right thing.
Frustration used to do the teaching
This part surprised me.
Nobody ever sat me down and taught me the first skill. A four-hundred-line function that I wrote at one in the morning and could not debug taught me. I felt that for two days. Doubting my own answers came from the compiler, which showed me my mistakes every day for a decade. Knowing where the checking stops came from launching something that looked fine. It was not fine, and a person I respected had to tell me.
None of that was mentoring. It was steady frustration over years, and the skills came out of it. We called it "experience" and treated it as something that happens to people. Nobody thought of it as something to teach.
Today, a developer four months into the job can build a working integration into a live client system in an afternoon. The typing is gone, and so are most of the small daily frustrations that used to teach these skills. The compiler still complains, but it complains to the agent now. The agent reads the error, fixes it and moves on before a person ever sees the message. The lesson still arrives every day. It arrives at the agent, which does not need it.
New developers are not taught less carefully than we were. They are taught just as carelessly. That worked when the job taught the lesson for free. It does not work now.
The five skills still matter. The daily frustration that taught them has gone.
So now someone has to teach them on purpose.
How to teach them on purpose
The first instinct is to write a checklist. We tried. Within a week, a checklist becomes the fortieth unread document in a repository.
The second instinct is a review gate, and we were tempted. It would not have caught this bug. A gate checks what it can measure, and everything measurable here passed. Every automated check we own would have called the wrong ticket type correct. That is the whole point of the story.
What has worked is smaller, and more tedious, than either.
Ask for the ticket back in their own words, before any code exists. Do not ask "do you understand?", because everyone says yes. Ask: "Say what this is asking for, and who it is for." If they cannot say it, you have found the problem before anyone builds on the misunderstanding. This costs nothing, and almost nobody does it.
Ask for a prediction before anything slow or risky runs. "What do you think will happen? Guess." Then compare out loud. Being wrong at this moment costs nothing, and it is the fastest way to find the gap in someone's understanding of the system. It is also the first step toward proving your own answer wrong, which is the hardest of the five skills to teach directly. Asking a developer four months in to state an architectural hypothesis only makes them feel behind. Asking them to guess what a command will print is the same skill, at a size they can manage.
Send them to look at the result where a real person would see it. Skip the terminal output and the agent's summary. Open the queue the way the client will open it. That check takes forty seconds, needs no expertise in the system, and is the only one that would have caught the broken tickets.
Make the agent ask the questions. I would not have predicted this one. We put the sequence into a skill file that runs on the junior developer's machine, written as instructions to the agent. The agent now asks her to explain the ticket back before it starts building. It asks what she expects before it runs something slow. It will not merge anything unless she approves that specific action, and an approval buried in a sentence about something else does not count. It tells her which layer it could not check and where she should go to look.
A document teaches only when someone reads it. The agent is present at the exact moment each of these questions matters, which no document and no senior developer can be.
How the agent asks turned out to matter more than what it asks, and most of that is restraint. It asks one question at a time, because a list of six is overwhelming. It asks, then waits, and does not answer its own question. It does not accept "yeah, makes sense." After two real attempts, it stops asking and tells her the answer, because working through a problem helps and being stuck does not. And she can say "just do it" at any time. The agent then does the work at once and mentions, once, what it skipped. People find a way around a teaching tool they cannot switch off, and they find it within a week.
The part you cannot hand to an agent
All five skills come down to the same job: deciding what correct means, then finding out whether you got it.
An agent cannot own that job. It will tell you, accurately and in good faith, that the interface accepted the data and the fields came back matching. Whether a bug report that arrives already resolved is useful to anyone is a question about the real world. It belongs to whoever understands what the feature is for.
That is the skill worth mentoring now. It never depended on writing code. We used to learn it as a side effect of the work.
Where this comes from
One junior developer, one afternoon, on a codebase we maintain for a client, plus two months of transcripts from the senior engineer on the same codebase. We have changed the details of the client system throughout.