← All articles

Product

Your agents can build anything. The Evidence Loop tells them what’s worth building.

The Dipio Evidence Loop runs the research that ships: real interviews, behaviour diagnostics, and evidence your coding agent builds from directly.

Julian UstiyanovychFounder, Dipio8 August 20267 min read
InterviewDiagnoseSpecShipLearn

The shift

Implementation got cheap. Knowing what to build didn’t.

Code that used to take a skilled engineer a week, a coding agent now writes in an afternoon. More than 80% of the code Anthropic merges into production is written by Claude today, up from low single digits when Claude Code launched at the start of 2025. What that has done to the people who used to write it manually:

In December 2025, Boris Cherny, the creator of Claude Code, confirmed that over the “previous thirty days”, 100% of his contributions to it were written by Claude Code. At the same time, in its AI-Assisted Engineering: Q1 2026 Impact Report, DX puts industry AI adoption at 93%. What reads as a frontier-lab quirk is already the industry default, and building has never been faster.

This month during YC Startup School, Garry Tan acknowledged the same pattern: “…the multiplier for coding is not just for coding. It’s for every piece of knowledge work. And it’s not just me. At YC, we get to watch this at portfolio scale. A year and a half ago, in the Winter ’25 batch, a quarter of the companies had codebases that were 95% AI-generated. Those companies use AI agents for everything now, not just code. And that batch is on track to becoming one of the fastest growing, most profitable batches in the history of YC.” The line that matters is the next one: “those companies use AI agents for everything now, not just code.”

At the same time, look at what all that speed is still measured against by the vast majority of organisations: (1) deployment frequency, (2) lead time, (3) review time, and (4) change failure rate. Every dashboard in engineering tracks how fast and how cleanly code ships. Not one of them tells you whether the thing was worth shipping. DORA, which defined those metrics, is blunt about where speed without direction leads: “In an era where AI allows us to build faster than ever, a user-centric focus acts as our steering wheel. Without it, we risk simply crashing faster.” (DORA, 2025)

Speed was never the bottleneck. Now that anyone can build almost anything, the only question that pays is whether it is worth building, and that is the one the usual toolkit cannot answer. Product analytics tells you what users did. Surveys tell you what they say. Neither tells you what is worth building next.

It is 2026, we have “almost” reached AGI, and we still bet on what to build the way we did in 2021.

The bottleneck

Validated intent is the bottleneck.

The missing piece has a name: validated intent. Proof of what your users actually need, specific enough for your team or your agent to build from.

The companies with the most sophisticated experimentation programs have already measured how often a shipped idea actually works. Across Microsoft’s experiments, only about 1/3 of tested ideas improved the metric they were designed to move. Pendo, examining real usage across 615 products, found that 80% of features are rarely or never used. At Booking.com, roughly 9 in 10 tested ideas fail to move their target metrics.

Teams do not build too slowly. They build the wrong things at scale. DORA calls it the “feature factory trap”: rewarding output, features shipped, over outcomes, the value a user actually gets. Its 2025 research found teams that resist it and hold a user-centric focus see 40% higher organisational performance. Coding agents do not correct this, they accelerate it. Without a way to know what is worth building, faster just means more waste.

Tan’s talk also has the frame for fixing it. Later in the same hour he lays out what he calls the equation for the next decade of your life: “a frontier model, which is rented and a commodity and getting cheaper by the quarter, + your context, which is owned by you and unique, and ideally nobody else on this earth has it, + a harness that wires them together… add that up and that gives you an agent that acts like a very fast version of you. Model quality is rented, but your brain is owned, ideally by you.” Marshall McLuhan called technology an extension of man, and Steve Jobs called the computer a bicycle for the mind. Tan’s update: “if you have what I’m describing here, then you have a self-driving rocket.”

Two of those three terms are already commodities. Everyone rents the same models, and the harnesses are converging: Claude Code, Codex, Cursor, all pointed at the same repository. Only the context is genuinely yours. Tan means a personal library of everything he has ever written. For a product team it means something rarer: evidence about what your users actually need. Nobody can rent that, and almost nobody is compounding it.

That is the knowledge work still waiting for its multiplier: deciding which code is worth writing.

Dipio Evidence Loop

Five steps, one continuous agentic workflow loop.

Alexandr Wang argued that the agentic loop itself is the thing worth building. Companies, he said, are just large-scale feedback loops where humans are operating each of the edges, and building agentic systems that can operate and optimise those loops is where the alpha sits.

Product development is one of those loops, and automating its edges is the easy part. Voice agents already run interviews at scale, and transcripts are cheap. What still decides the outcome is which questions get asked, and whether anyone can explain the why underneath the answers. Ask the wrong things and you get fluent, well-transcribed confirmation of something that never mattered, and then you build it faster than ever.

We have a view about this, and it is not a neutral one. We are behavioural scientists. Our research is in machine psychology, the field that turns the experimental methods of human psychology on language models themselves: what it takes for a model to stand in for a real person, and how to reach the actual driver of a decision rather than the reason someone gives you when you ask. People are unreliable witnesses to their own behaviour. No amount of interview volume fixes that. Handling it is what the science is for.

So the loop we argue for has rules. It has to be able to say no: an agent that never tells you the evidence is missing is not doing research for you, it is agreeing with you faster. Evidence has to arrive before the build, not after it, or it is a justification. The diagnosis has to explain why people behave as they do, not just report what they said, because a theme is not a cause and you cannot design against a theme. And what comes out has to be specific enough for an agent to act on without a human rewriting it in the middle, or the loop quietly becomes a workflow again.

01INTERVIEW

An AI interviewer runs deep, structured voice conversations with your real users, no moderator needed.

02DIAGNOSE

The diagnosis pinpoints what blocks your users and what to build next, grounded in psychological and behavioural science.

03SPEC

The diagnosis becomes an implementation-ready spec, grounded in traceable evidence and ready for your coding agents.

04SHIP

Your coding agents and engineers build straight from the spec, with no translation step in between.

05LEARN

Every release ships as a measurable experiment, and the results become the next question to study.

the next question loops back to Interview

The science sits under the first two steps: what gets asked in the interview, and what gets made of the answers. It matters most where the users are synthetic, because a twin is only worth talking to if the model behind it behaves like the person it stands for. That work is independently validated rather than asserted, and our first paper is in press at Behavioural Public Policy (Cambridge University Press). Read the paper

What we are building

Evidence-based machinery, and the science under it.

That is the work: the machinery that makes this loop run, and the behavioural science that decides what counts as evidence in the first place. One without the other is either a pipeline moving unverified opinion at speed, or a good study nobody builds from.

The future we want is not complicated. Teams ship fast. What they ship is personalised, because they know who they are building for. And it addresses a pain the customer actually has, rather than one somebody assumed on their behalf.

Build from evidence.

Connect your coding agent to the Dipio Evidence Gate, or talk to us about running the loop end to end.