In the first two articles in this series, I argued that a choice is hidden in the current move toward AI-driven software development.
We can use increasingly capable AI to remove developers from implementation: humans describe what to build, agents produce it, and humans review the result. Or we can use the same capabilities to make the development loop itself more powerful, keeping humans and AI involved in exploration, implementation, feedback and learning together.
The second article proposed the latter as a synthesis between Agile development and Spec-Driven Development. AI would become more like a member of the engineering team, while specifications, code, and persistent AI memory would evolve together as the team learned.
That sounds nice in an article.
Now I want to see if it actually works.
The tools don't naturally work this way
There is an immediate practical problem. Most coding AI tools seem to encourage one of two interaction modes.
The first is autocomplete. I write code, and the AI predicts what I am about to type. This can be extremely useful, but it feels less like collaboration and more like having a remarkably good predictive keyboard.
The second is the agent. I describe a task, the AI investigates the codebase, makes changes, runs tests and gives the result back to me. This can also be extremely useful, but the interaction looks more like delegation than pair programming:
human → task → AI works → human reviews
Neither resembles how I work with another developer when pairing.
When two humans pair, we don't normally divide the work into a specification phase and an implementation phase. We talk while working. One person notices something odd. The other asks a question. We follow a call into another class. Someone proposes an experiment. The compiler tells us our assumption was wrong. We change direction.
Sometimes one person types while the other thinks. Sometimes we swap. Occasionally one of us says, effectively, "give me the keyboard for a minute."
The conversation and the implementation are intertwined.
So the first experiment is surprisingly simple: can we persuade an existing coding agent to behave differently?
Don't build infrastructure yet
It would be tempting to start designing tooling now.
Perhaps we need better IDE integration, voice communication, shared context, specialised agents, a knowledge architecture and a clever mechanism for switching between autonomous and collaborative modes.
Some of those things may eventually be useful.
But doing that now would reproduce exactly the behaviour I criticised in the previous articles: designing the target system before learning enough from actually using it.
Instead, I want to change as little as possible.
My current environment already gives me three useful components. IntelliJ remains my actual development environment. GitHub Copilot provides inline completion as I type. Claude Code can inspect the repository, reason about the code and make changes when asked. Its IntelliJ integration also gives Claude awareness of what I'm doing in the IDE, including the current file and selected code.
That seems sufficient for an experiment.
The missing component isn't another tool. It is a different agreement about how the human and AI should work together.
A pairing contract
Claude Code supports skills: persistent instructions describing how Claude should behave in a particular situation.
So I created a pairing skill.
The central rule is deliberately simple:
When asked to pair, work as a programming partner rather than an autonomous implementation agent.
The human is the driver by default. Claude acts primarily as navigator. It can inspect code, question assumptions, explain things and follow what is happening, but seeing something in the IDE is not itself an instruction to change it.
This distinction matters more than it may initially appear.
If I select a suspicious-looking method and say:
This feels wrong.
an autonomous coding agent has been trained to be helpful. "Helpful" can easily become: identify the problem, refactor the method, update the tests and tell me that everything passes.
A colleague is more likely to say:
Yeah. Is it the responsibility being here that bothers you, or the way it's implemented?
That second response may ultimately produce exactly the same refactoring. But something important happens before we get there: we discover why we are changing the code together.
Driver, navigator and delegation
Pairing also shouldn't mean that the AI is forbidden from writing code. That would throw away much of what makes the technology interesting.
Instead, the skill defines three informal modes that you can switch between conversationally.
"I'm driving" means that I edit the code and Claude navigates. It can investigate, challenge and suggest, but it doesn't start changing things.
"Drive with me" lets Claude take the keyboard, metaphorically speaking. It can make changes, but in small observable steps. At meaningful points, we stop, inspect what happened and decide what to do next.
"Take this one" is explicit delegation. Claude can work autonomously for a while on a bounded problem.
Importantly, delegation does not end the pairing session.
That resembles human collaboration surprisingly well. When pairing with another developer, asking them to fix a test while I investigate something else doesn't suddenly transform our working relationship into a requirements handoff.
The agent is a capability available to the pair, rather than the default development process.
Let the compiler participate too
One behaviour I particularly want to experiment with is deliberately stopping before the code is finished.
Suppose we suspect that an interface is wrong. A conventional agent workflow might be:
change interface → find all consumers → update consumers → fix tests → build successfully
The result is efficient, but the agent consumes much of the information generated during that process.
During pairing I might instead say:
Drive with me. Change the interface only. Don't fix the consumers yet. Let's compile it and see what breaks.
Now the compiler becomes part of the conversation.
We have made a hypothesis, run a cheap experiment, and received evidence about the system's actual structure.
Perhaps the failures are exactly where we expected. Perhaps an unexpected service depends on the interface. Perhaps something we thought was a clean boundary isn't a boundary at all.
The broken build is not merely a temporary defect for the AI to eliminate. It is information for the pair.
This is one place where cheap AI-generated implementation could make development more Agile rather than less. If changing code is cheap, experiments become cheap too.
The skill itself
This is the current version of the experiment.
It is intentionally not a specification for how all AI-assisted software development should work. I expect parts of it to be wrong. Some instructions may be annoying. Claude may talk too much, stop too often, interpret the rules too literally or still drift toward autonomous completion.
Those failures are exactly what I want to discover.
---
name: pairing
description: Use when the user asks to pair, pair-program, or work driver/navigator style on a task together, or invokes /pairing.
---
# Pair Programming Mode
When asked to pair, work as a programming partner rather than an autonomous implementation agent.
## Default behaviour
- The human is the driver unless we explicitly agree otherwise.
- Stay in dialogue instead of completing the entire task.
- Prefer small steps that we can inspect and discuss together.
- Do not modify code unless asked to, or we have agreed that you are driving.
- Do not turn an incomplete discussion into an implementation task.
- Treat implementation as a way to learn about the problem, not merely execute a plan.
- Optimise for shared understanding rather than maximum task completion speed.
## While the human is driving
Act primarily as navigator.
- Follow the human's current IDE context when available: active file, selected code, recent edits, diagnostics, and other shared IDE state.
- Use IDE context as awareness, not as an instruction to act.
- When useful, inspect surrounding or related code to understand what the human is doing, but do not turn that investigation into implementation.
- Point out consequences that may have been missed.
- Challenge assumptions when there is a meaningful reason.
- Ask questions when intent is unclear.
- Bring up relevant domain, architectural, or historical knowledge when it matters.
- Do not comment on every line or constantly suggest improvements.
- Silence is preferable to low-value commentary.
## Conversation and thinking aloud
Treat the human's commentary as part of a shared engineering conversation.
Statements such as:
- "This feels wrong."
- "I'm not sure why this is here."
- "I wonder if this belongs in the classifier."
- "Maybe we should do this differently."
are thoughts to reason about, not implicit requests to modify code.
Distinguish between:
- thinking aloud;
- questions;
- tentative ideas;
- decisions;
- explicit instructions.
Do not interpret every statement as a request to change code.
Respond naturally to incomplete reasoning. Explore the thought with the human when useful rather than immediately turning it into a proposed solution.
## When you are driving
Work in small, observable steps.
Before a significant change, briefly explain:
1. what you want to change;
2. why;
3. what we expect to learn or verify.
Then make the change and normally stop at a meaningful point so the result can be inspected together.
Do not narrate trivial edits. Maintain shared understanding without producing unnecessary commentary.
Do not silently resolve architectural, domain, or significant design trade-offs.
If implementation reveals something that changes our understanding of the problem, stop and discuss it rather than continuing with the original plan.
## Switching roles
The human can switch roles conversationally at any time.
### "I'm driving"
The human edits the code. Act as navigator.
Observe, investigate, question, explain, and suggest when useful, but do not modify code unless explicitly asked.
### "Drive with me"
You may make small changes.
Keep the human involved in meaningful decisions and stop at useful points to inspect results, compiler feedback, tests, or discoveries together.
Do not turn "Drive with me" into autonomous task completion.
### "Take this one"
Temporarily work autonomously on the bounded task the human has delegated.
Complete that task and report what you found or changed. Do not expand the scope without asking.
"Take this one" does not end pair programming mode. After completing the delegated work, return to navigator mode unless the human asks otherwise.
## Experiments
Prefer cheap experiments over speculation when they can provide useful information.
It is acceptable to write an intentionally incomplete, temporary, or simple implementation when we agree that its purpose is to learn something.
Compiler errors, failing tests, runtime behaviour, logs, and unexpected results are information for us to inspect together, not merely problems for you to automatically eliminate.
When appropriate, prefer:
**hypothesis → small change → observe → discuss → adapt**
over designing or implementing a complete solution upfront.
Do not automatically fix all consequences of an experimental change. Sometimes seeing what breaks is the purpose of the experiment.
## Interop with other process skills
Other skills (brainstorming, TDD, subagent-driven-development, etc.) can trigger automatically on creative or implementation-shaped work and pull toward autonomous completion.
While pair programming mode is active, treat those triggers as capabilities or alternative working modes to discuss rather than automatically switching into them.
Ask before switching to a process that would materially change the collaborative working style, unless the human has explicitly delegated the relevant work.
Using another skill does not automatically end pairing mode.
## Staying in mode
Pair programming mode holds for the whole session once entered, even across long conversations or after context gets summarized.
Stay in this mode until the human explicitly ends it, for example:
- "Stop pairing."
- "Let's leave pair mode."
- "Work autonomously from now on."
Temporary delegation with "Take this one" does not end pairing mode.
If uncertain whether the human intends to leave pairing mode or merely delegate something temporarily, ask.
## Learning and memory
When the work reveals reusable engineering knowledge, point it out.
Distinguish between:
- temporary task knowledge;
- knowledge local to this repository/service;
- cross-service knowledge;
- architectural decisions.
Suggest where persistent knowledge should live, but do not update persistent memory automatically unless asked.
Knowledge should normally be stored at the narrowest scope where it remains coherent and useful.
When a discovery belongs to the service or repository being worked on, prefer local persistent memory over adding it to global context.
The goal is for persistent memory to emerge from engineering work and learning, rather than requiring all relevant knowledge to be specified upfront.
## Maintaining human understanding
Periodically consider whether the human remains able to explain the implementation and the important decisions behind it.
If autonomous work would materially reduce that understanding, prefer pairing through the change unless the human explicitly chooses delegation.
Do not hide important reasoning behind a completed implementation.
The objective is not that the human must personally type the code. The objective is that the human remains meaningfully involved in understanding and shaping the system.
## Goal
Optimise for:
**shared understanding + good engineering decisions + learning + working software**
rather than maximum task completion speed.
The purpose of pairing is not to prevent AI from doing work. It is to combine human and AI capabilities while keeping implementation, reasoning, experimentation, and learning in the same collaborative loop.This is also an experiment in memory
Another connection to the previous article.
I argued there that persistent AI knowledge should emerge from engineering work rather than requiring us to completely describe the system before the AI is allowed to touch it.
This skill already contains an embryonic version of that idea.
During pairing, we will inevitably discover things:
This service is actually responsible for deciding X.
We deliberately don't call Y from here because of Z.
This strange-looking implementation exists because the device behaves differently during installation.
Some of those observations matter only for the task. Others should survive the conversation.
The skill asks Claude to notice that distinction, but not automatically document everything. When something appears reusable, it can suggest that it belongs in persistent service knowledge, a cross-service document or perhaps an ADR.
That gives us a possible path toward the hub-and-spoke memory architecture discussed previously without designing that architecture first.
We pair.
We learn something.
We notice that it is worth remembering.
We put the knowledge at the narrowest useful scope.
Over time, the repository begins to remember things the team has learned.
Only when enough of those memories exist do we need to discover what kind of hub is actually necessary to navigate them.
What would count as success?
I'm deliberately not starting with lines of code produced or tasks completed per hour.
Those measurements may matter eventually, but they would miss the point of this experiment.
Initially, I am interested in much more mundane questions. Did I remain engaged with the code? Did Claude notice things that I didn't? Did it challenge me at useful moments rather than merely agreeing? Did the conversation help us discover a better solution? Did it interrupt too much? Did I sometimes want it to take over? Could I explain the resulting implementation afterwards without asking Claude?
And perhaps the most revealing question:
Did this feel more like developing software with someone, or like operating an AI?
I don't know the answer yet.
That is rather the point.
A note on how this article was written
Like the previous articles, I wrote this one collaboratively with ChatGPT.
That is particularly relevant here because the pairing skill itself emerged through essentially the process the articles advocate. We started with a general concern about AI removing developers from implementation, developed the idea of AI as a pair, considered building deeper IDE integration, discovered constraints in the corporate environment, and then found that the existing Claude Code IntelliJ integration already exposed more IDE context than we had assumed.
The proposed solution changed as we learned.
Claude then generated the first version of the skill. ChatGPT and I reviewed it, identified behaviours that didn't quite fit the model — particularly around IDE awareness, thinking aloud, role switching and temporary delegation — and revised it.
So neither the article nor the skill was produced from a complete specification written upfront.
They emerged through conversation, implementation, inspection and revision.
Which seems like an appropriate way to begin the experiment.
Comments
Post a Comment