You don't watch a colleague type. Why watch the machine?
Marius Wehrle, Katarina Kostic ·

Out of the chat

Chat is fine. Chat is staying. Asking a question in plain words and getting a straight answer is one of the nicer things to happen to computers.

But the real gain from AI isn't a better chat. It's the work that happens while you're not in the room. Software got there first, so that's where it's easiest to see.

Seconds, minutes, hours #

First it was seconds. A coding assistant suggested the next line before you'd finished typing. Lovely. Nobody called that waiting.

Then it was minutes, which is where most of us are now. You half-watch. You check your phone. You'd never stand behind a colleague like this. Two meters are running, like in a taxi: the machine's and yours, and only one of them is cheap. Eventually you open a second window and start a second job.

Now it's hours, sometimes days, and nobody sits through that. That's the opportunity.

Why you stop checking #

Waiting has a second cost, and this one is the risky one.

A three-second suggestion gets a glance. Fair enough. It's one line, and if it's wrong you'll notice by lunch.

A five-minute answer gets… also a glance, not because it's small but because you've been sitting there for five minutes and checking would mean sitting there longer. So out come the magic words: go ahead, sounds good, what would you suggest? You tell yourself that's oversight. It's agreeing with extra steps.

Now picture six hours of work that happened while you were at lunch, in a meeting, and briefly asleep. Half an hour to check it? Bargain. And you're actually good at it now: rested, unbothered, no sunk afternoon to defend, reading it the way you'd read a report from a colleague.

Which gives you the odd rule of this whole business: the longer the job, the more checking it earns and the better you are at it. Long jobs aren't the scary ones. They're the ones you finally read properly.

We did exactly this last week. For three days, agents in our own harness and in external ones worked side by side through the research behind our new hybrid search: more than 200 combinations of keyword, meaning and link search, tested across different benchmarks and setups, and every one written up. Nobody watched. Then we read it the way you'd read a colleague's report, properly.

And none of this is really about software. The same shape fits any work that takes hours, repeats every week, or keeps an eye on something until it changes: the monthly figures that take a day to pull together, the contract review nobody gets round to, the market you'd like watched every week.

What it actually takes #

The computer didn't earn its place by writing letters faster. AI won't earn its place by answering chat questions faster either. Faster is a percentage. Not watching is a multiplier: a percentage is a nicer afternoon, a multiplier is the two-day job that happened while you were asleep. Getting there isn't a model problem. It's the plumbing around it: the infrastructure that decides what runs and where the results go.

You need a way to control what runs: what it may see, change and spend, and when it has to stop and ask. You need somewhere for results to land that isn't a tab you'll close, and that somewhere has to be yours. What comes out of these jobs is valuable. It belongs to the organisation, not to whichever laptop it ran on, and it needs the kind of shared, versioned control that software teams have had for years. No lock-in. Your results, your rules.

The fair objection: six unwatched hours can also be six hours of a wrong assumption compounding. That risk is real. It's why the controls come first, and why the review at the end is a proper read, not a glance.

Internally we've been running on something we call Continuous Inquiry: a way of working on a problem over days rather than minutes, with what it learns saved somewhere it can be found again, and shared across people and tools instead of dying in a tab. It runs in a workspace we had to build ourselves, where people work with each other and with agents under one set of rules, whichever tool an agent runs in. That's a post of its own. With luck, by the time we write it we'll also have cracked the other problem: guessing how long any of this takes before it starts. Currently: somewhere between ten minutes and Thursday.

Marius Wehrle, PhD and Katarina Kostic

Human thinking. AI typing. Human signing.

Share this post