Engineering practice

The parts of programming that did not get automated

An agent filed three tickets into a client's system, read every field back, and confirmed that each one was correct. The feature was still broken. The missing skill had nothing to do with writing code. It is one of five skills that almost nobody teaches on purpose, because for thirty years developers picked them up on their own.

By Ashish Tonse 10 min read

Skills that still matter
5

A developer on our team spent an afternoon getting a bug-report form to file tickets into a client's ticketing system. It worked. The agent filed three tickets, then queried the system and read every field back: title, description, steps to reproduce, and the reproducibility dropdown. Each field was correct. The agent reported success, and the report was true.

The feature was still broken.

The form created tickets of the wrong type. The workflow for that type marks an item resolved the moment it arrives, so every new bug report arrived already closed, where nobody filtering for open work would see it. The agent had noticed the symptom. It wrote that a newly filed bug in a resolved status would not appear in open-work filters. Then it blamed the client's configuration and suggested raising it with them.

The client's configuration was fine. The issue type was set in our code.

Every new bug report arrived already resolved The bug-report form, in our code, files tickets into the client's ticketing system. The system has an open column, which is what the filters for open work show, and a resolved column. The form set the wrong issue type, and the workflow for that type marks an item resolved the moment it arrives. So each of the three tickets passes the open column and lands straight in resolved, and the open column shows no new bug reports. The agent read every field back and every field was correct. The issue type was set in our code, not in the client's configuration. OUR CODE Bug-report form SETS THE WRONG ISSUE TYPE THE CLIENT'S TICKETING SYSTEM Open WHAT THE OPEN-WORK FILTERS SHOW No new bug reports Resolved SET ON ARRIVAL BY THE WORKFLOW Ticket 1 FIELDS CORRECT Ticket 2 FIELDS CORRECT Ticket 3 FIELDS CORRECT THE AGENT BLAMED THE CLIENT'S CONFIGURATION THE ISSUE TYPE WAS SET IN OUR CODE
Each ticket arrived with every field correct and landed straight in Resolved, so the view of open work stayed empty. The wrong issue type came from our form, not from the client's system.

Nothing here was a coding problem

Every line the agent wrote was correct. Every check it ran passed, and the passes were real. If you had graded that session on what we have graded developers on for decades, whether the code works, it would have scored full marks.

What was missing was someone who doubted that "it works" was the whole answer.

That kind of doubt made a good engineer in 1998, and it matters just as much now that a machine does the typing. The list of things that make someone good at this job has barely changed. What has changed is that the skill we always taught first and hired for, writing code that works, is now the one an agent does well. Whole interview formats exist to test that one skill, and companies still use them.

Five skills did not get automated. "Critical thinking" is too broad to teach, but these five can be taught. The examples below are all ours: one junior developer's afternoon, and two months of transcripts from the senior engineer on the same codebase.

The five

Breaking a problem into parts that can each be wrong on their own. The goal is that when something is wrong, you can tell which part. In our transcripts, the senior version of this is an instruction sent before any work started: "Do phase 1 only: bump every place the version is pinned, and tell me what a v2 migration would cost here before doing it." Two pieces, one of them held back, with a question attached to the one held back. Nothing stops an agent from working this way. Someone has to ask it to.

Trying to prove your own answer wrong. The strongest example in two months of transcripts is eleven words, typed at three in the morning and never edited:

you can check if it does. i doubt it. but check.

That message holds two instructions: test this, and here is what I already believe. Ask an agent "how should we do this?" and it produces something reasonable, then agrees with itself. Give an agent a belief and tell it to test it, and it has something to push against. The output is completely different, because a person held an opinion first.

Asking for an answer, or giving a belief to test Two ways to prompt an agent. On the left, ask it "how should we do this?" and it produces something reasonable, then agrees with itself, because it has nothing to push against. On the right, the eleven-word message typed at three in the morning: "you can check if it does. i doubt it. but check." It holds two instructions. "You can check if it does, but check" means test this. "I doubt it" says what the person already believes. The agent now has something to push against, and the output is completely different. ASK FOR AN ANSWER "how should we do this?" Something reasonable THE AGENT'S FIRST ANSWER THEN AGREES WITH ITSELF NOTHING TO PUSH AGAINST GIVE IT A BELIEF TO TEST ELEVEN WORDS, 3 A.M. "you can check if it does. i doubt it. but check." "check if it does ... but check" INSTRUCTION 1: TEST THIS "i doubt it" INSTRUCTION 2: WHAT I ALREADY BELIEVE Something to push against A DIFFERENT OUTPUT
An open question gets a reasonable answer that the agent then agrees with. A stated belief plus an order to test it gives the agent something to push against.

Not accepting the first plausible answer. This reflex is separate from the last one. It comes when the answer arrives, where the last one comes before. "what? that makes no sense, every sprint before this, we've been squash merging and all the details were in the PR description. why would we change that now? explain please." The person asking did not know the answer. They noticed that it contradicted something they knew, and said so before moving on.

Saying what you expect before the work starts. Our best example of this came from the junior developer. Before touching anything, she wrote: "all i need to verify is that on localhost i go to report a bug and submit and it gets submitted in jira." That is a clear acceptance test, written in advance, in plain language. She already had this skill.

Knowing where the checking stops. An agent checks the layer it can reach. In the ticket story, that layer was whether the interface accepted the data and whether the values arrived intact. Both were true, and both were the wrong question. Whether a bug report is the right kind of ticket to create at all is a decision about what the feature is for, and reading the response back will never reveal it. The limit moves with the task. An agent that can run the test suite will confirm that the tests pass. That says nothing about whether the tests check the right thing.

What the agent checked, and what it could not Three layers of the bug-report feature. The agent checked that the system accepted the data, with three tickets filed, and that every field came back intact: title, description, steps to reproduce and reproducibility. Both checks passed. It could not check whether a bug report was the right kind of ticket at all, because that depends on what the feature is for. It was the wrong kind: each new report arrived already resolved and was hidden from filters for open work. Opening the queue the way the client will, which takes about forty seconds, would have caught it. Did the system accept the data? THREE TICKETS FILED CHECKED, PASSED Did every field come back intact? TITLE, DESCRIPTION, STEPS, REPRODUCIBILITY CHECKED, PASSED THE LAYER THE AGENT CAN REACH WHAT THE FEATURE IS FOR Is a bug report the right kind of ticket? IT WAS NOT: EACH ONE ARRIVED ALREADY RESOLVED HIDDEN FROM FILTERS FOR OPEN WORK NOT CHECKED READING THE FIELDS BACK WILL NEVER SHOW IT THE CHECK THAT WOULD HAVE CAUGHT IT: OPEN THE QUEUE THE WAY THE CLIENT WILL. ABOUT FORTY SECONDS.
Everything the agent could reach passed. The problem sat in the layer above, which depends on what the feature is for. Opening the queue the way the client will would have caught it in about forty seconds.

Frustration used to do the teaching

This part surprised me.

Nobody ever sat me down and taught me the first skill. A four-hundred-line function that I wrote at one in the morning and could not debug taught me. I felt that for two days. Doubting my own answers came from the compiler, which showed me my mistakes every day for a decade. Knowing where the checking stops came from launching something that looked fine. It was not fine, and a person I respected had to tell me.

None of that was mentoring. It was steady frustration over years, and the skills came out of it. We called it "experience" and treated it as something that happens to people. Nobody thought of it as something to teach.

What used to do the teaching Three frustrations and the skill each one left behind. A four-hundred-line function, written at one in the morning, that could not be debugged and was felt for two days, taught breaking a problem into parts. The compiler, which showed every mistake every day for a decade, taught doubting your own answers. A launch that looked fine but was not, which someone respected had to point out, taught knowing where the checking stops. None of it was mentoring. It was called experience. WHAT DID THE TEACHING THE SKILL IT LEFT BEHIND A four-hundred-line function, at one in the morning IMPOSSIBLE TO DEBUG, AND FELT FOR TWO DAYS Breaking a problem into parts The compiler SHOWED EVERY MISTAKE, EVERY DAY, FOR A DECADE Doubting your own answers A launch that looked fine IT WAS NOT, AND SOMEONE RESPECTED HAD TO SAY SO Knowing where the checking stops NONE OF IT WAS MENTORING. WE CALLED IT EXPERIENCE.
Three of the five skills, and the everyday frustration that taught each one. Nobody planned any of it.

Today, a developer four months into the job can build a working integration into a live client system in an afternoon. The typing is gone, and so are most of the small daily frustrations that used to teach these skills. The compiler still complains, but it complains to the agent now. The agent reads the error, fixes it and moves on before a person ever sees the message. The lesson still arrives every day. It arrives at the agent, which does not need it.

The lesson now arrives at the agent Where the compiler's error message goes, then and now. Then, the error went straight to the developer, who read it and learned from it. Now, the error goes to the agent, which reads it, fixes it and moves on, and the developer never sees the message. The lesson still arrives every day. It arrives at the agent, which does not need it. THEN The compiler AN ERROR MESSAGE The developer READS IT, AND LEARNS FROM IT ERROR NOW The compiler THE SAME ERROR MESSAGE The agent READS IT, FIXES IT, MOVES ON The developer NEVER SEES THE MESSAGE ERROR THE LESSON STILL ARRIVES EVERY DAY. IT ARRIVES AT THE AGENT, WHICH DOES NOT NEED IT.
The compiler still complains every day. The agent reads the error, fixes it and moves on, so the developer never sees the lesson.

New developers are not taught less carefully than we were. They are taught just as carelessly. That worked when the job taught the lesson for free. It does not work now.

The five skills still matter. The daily frustration that taught them has gone.

So now someone has to teach them on purpose.

How to teach them on purpose

The first instinct is to write a checklist. We tried. Within a week, a checklist becomes the fortieth unread document in a repository.

The second instinct is a review gate, and we were tempted. It would not have caught this bug. A gate checks what it can measure, and everything measurable here passed. Every automated check we own would have called the wrong ticket type correct. That is the whole point of the story.

What has worked is smaller, and more tedious, than either.

Ask for the ticket back in their own words, before any code exists. Do not ask "do you understand?", because everyone says yes. Ask: "Say what this is asking for, and who it is for." If they cannot say it, you have found the problem before anyone builds on the misunderstanding. This costs nothing, and almost nobody does it.

Ask for a prediction before anything slow or risky runs. "What do you think will happen? Guess." Then compare out loud. Being wrong at this moment costs nothing, and it is the fastest way to find the gap in someone's understanding of the system. It is also the first step toward proving your own answer wrong, which is the hardest of the five skills to teach directly. Asking a developer four months in to state an architectural hypothesis only makes them feel behind. Asking them to guess what a command will print is the same skill, at a size they can manage.

Send them to look at the result where a real person would see it. Skip the terminal output and the agent's summary. Open the queue the way the client will open it. That check takes forty seconds, needs no expertise in the system, and is the only one that would have caught the broken tickets.

Make the agent ask the questions. I would not have predicted this one. We put the sequence into a skill file that runs on the junior developer's machine, written as instructions to the agent. The agent now asks her to explain the ticket back before it starts building. It asks what she expects before it runs something slow. It will not merge anything unless she approves that specific action, and an approval buried in a sentence about something else does not count. It tells her which layer it could not check and where she should go to look.

A document teaches only when someone reads it. The agent is present at the exact moment each of these questions matters, which no document and no senior developer can be.

How the agent asks turned out to matter more than what it asks, and most of that is restraint. It asks one question at a time, because a list of six is overwhelming. It asks, then waits, and does not answer its own question. It does not accept "yeah, makes sense." After two real attempts, it stops asking and tells her the answer, because working through a problem helps and being stuck does not. And she can say "just do it" at any time. The agent then does the work at once and mentions, once, what it skipped. People find a way around a teaching tool they cannot switch off, and they find it within a week.

Where the agent asks, and how Four points in one ticket where the agent asks the junior developer a question. Before any code, it asks her to explain the ticket back: what it asks for and who it is for. Before a slow run, it asks her to guess what will happen, then compares out loud. After it runs, it names the layer it could not check and where she should look. Before a merge, it needs her approval for that specific action, and a yes inside another sentence does not count. How it asks: one question at a time, it waits for her answer, it gives the answer after two real attempts, and when she says just do it, it does the work and says once what it skipped. WHERE THE AGENT ASKS, DURING ONE TICKET BEFORE ANY CODE Explain the ticket back WHAT IT ASKS FOR AND WHO IT IS FOR BEFORE A SLOW RUN Guess what will happen THEN COMPARE OUT LOUD AFTER IT RUNS The layer it could not check AND WHERE SHE SHOULD LOOK BEFORE A MERGE Approve this one action A YES INSIDE ANOTHER SENTENCE DOES NOT COUNT HOW IT ASKS One question at a time A LIST OF SIX OVERWHELMS Waits for her answer NEVER ANSWERS ITSELF Two real attempts THEN IT GIVES THE ANSWER "Just do it" works IT SAYS ONCE WHAT IT SKIPPED
The agent asks each question at the point in the ticket where it matters. How it asks, along the bottom, turned out to matter more than what it asks.

The part you cannot hand to an agent

All five skills come down to the same job: deciding what correct means, then finding out whether you got it.

An agent cannot own that job. It will tell you, accurately and in good faith, that the interface accepted the data and the fields came back matching. Whether a bug report that arrives already resolved is useful to anyone is a question about the real world. It belongs to whoever understands what the feature is for.

Who owns each part of the job The job all five skills come down to, in three parts. First, the person decides what correct means: what the feature is for and who it is for. Second, the agent builds it and checks what it can reach, such as whether the interface accepted the data and the fields came back matching. Third, the person finds out whether they got it, by looking where a real person will see the result. The first and third parts belong to whoever understands what the feature is for, and cannot be handed to an agent. THE PERSON 1 Decide what correct means WHAT THE FEATURE IS FOR, AND WHO IT IS FOR THE AGENT 2 Build it, check what it can reach THE INTERFACE ACCEPTED THE DATA, THE FIELDS CAME BACK MATCHING THE PERSON 3 Find out whether you got it LOOK WHERE A REAL PERSON WILL SEE THE RESULT 1 AND 3 CANNOT BE HANDED TO AN AGENT WHETHER A BUG REPORT THAT ARRIVES ALREADY RESOLVED IS USEFUL TO ANYONE IS A QUESTION ABOUT THE REAL WORLD.
The agent can build and check what it can reach. Deciding what correct means, and finding out whether you got it, stay with the person who knows what the feature is for.

That is the skill worth mentoring now. It never depended on writing code. We used to learn it as a side effect of the work.

Where this comes from

One junior developer, one afternoon, on a codebase we maintain for a client, plus two months of transcripts from the senior engineer on the same codebase. We have changed the details of the client system throughout.

Putting AI in front of people who are not engineers?

We are happy to talk it through. Bring the workflow and the questions your team has about it.

How we implement AI

Book a 30-minute call