Note
Keep Shooting
Everyone reads the last three years of AI as proof the machine can one-shot anything. Read it closer: every leap came from iteration. The one-shot was never real.
Open any timeline the week a new model drops and you get the same word, over and over: one-shotted. One-shotted a landing page. One-shotted a Three.js game. One-shotted a whole app, one prompt, no edits. The word carries a quiet claim underneath it: the machine is so good now that it hits on the first try. No aiming. No missing. No second shot.
I have spent the last three years building on these systems and watching them get better, and I read that story backwards from almost everyone. The one-shot is not what happened. It is the marketing. Underneath it, the entire arc of modern AI is one long argument for iteration. The machines people hold up as proof you can skip the grind are, every one of them, monuments to the grind.
Let me show you. Then I will tell you why it matters far away from any keyboard.
The models stopped trying to answer in one shot
Start with the models themselves.
For a while, getting more out of a model meant one thing: make it bigger. More parameters, more data, more pretraining compute. GPT-4 landed in March 2023 and that was the game. You wrote a clever prompt, the model did one forward pass, and whatever came out was your answer. One shot, by design.
Then the frontier did something telling. In September 2024, OpenAI shipped o1, a model "trained to generate long chains of thought before returning a final answer." In plain words: they taught it to stop answering immediately. To think, wander, check itself, and revise, privately, before it says a word. OpenAI's own framing was that accuracy climbs with "the amount of compute spent thinking before answering." A whole new way to make a model better, and it was not a bigger brain. It was giving the model room to iterate before it commits.
Within six months, everyone had followed. Google shipped a thinking Gemini that December. Anthropic made reasoning a built-in mode of Claude in February 2025. By the big flagship releases later that year, "think longer" was not a separate product. It was a dial inside the main model.
Sit with what that means. The single most important model advance of the last three years was teaching machines to not one-shot. We spent extraordinary money to buy them the exact thing a one-shot skips: another pass. A second look. Iteration, moved right into the silicon.
And that is only the visible half. The invisible half, the training, was always a loop. These models are shaped by feedback: generate, get scored, adjust, repeat, at enormous scale. That is what RLHF is, and what Anthropic's version is, where an AI critic grades outputs against written principles. Every point release is then tuned against a wall of evals that act as the feedback signal. Nobody one-shots a frontier model. They iterate one into existence, and then they teach it to iterate.
The agents that worked are loops, not prompts
Now the agents, because this is where the fantasy showed its face and lost.
March 2023: AutoGPT. You gave it a goal in one prompt and walked away, and it was supposed to just do the thing. Run the research, build the business, book the trip, autonomously, start to finish. It shot to the top of GitHub's trending charts in days. It was the purest one-shot dream we ever built: state the target once, get the whole result.
It did not work. The agents looped on themselves, forgot what they were doing, burned money down rabbit holes, and had no moment where a person could look at a missed shot and correct the aim. The most honest verdict came from the project itself. AutoGPT eventually abandoned open-ended autonomy and became a builder where humans design the boundaries and the AI executes inside them.
What replaced the fantasy was humbler, and it won. By the end of 2024 the working definition of an agent had settled, and Anthropic wrote it down cleanly: agents are "typically just LLMs using tools based on environmental feedback in a loop." Read that again. Not one prompt. A loop. Act, look at what the world says back, adjust, act again. The exact opposite of a one-shot.
You can watch the payoff on the scoreboard. SWE-bench asks a system to fix real bugs in real codebases. When it launched in 2023, the best setups solved under 2% of the tasks. By late 2024, agents with proper feedback loops were around half. By mid-2025 the top numbers were in the seventies. That climb did not come from a model that finally one-shot the answer. It came from harnesses that let a model try, run the tests, read the failure, and go again.
Here is the part that should end the argument. The frontier of autonomy today is measured in how long an agent can work unattended. The research group METR tracks it, and the length of task an agent can finish on its own has been roughly doubling every few months, from seconds a couple of years ago to many hours now. That sounds like the machine finally one-shotting big things. It is the reverse. A sixteen-hour autonomous run is not one shot. It is thousands of shots, fast, each one aimed by the result of the last. Nobody found a way around iteration. They just made it cheap enough to do at a speed you cannot match by hand.
"Give it more tries" is a measurable superpower
Even the way we score these systems admits it.
There is a metric that has run quietly underneath this whole era, from a 2021 paper: pass@k. pass@1 asks, did the very first answer work. pass@k asks, did any of k answers work. The gap between them is the entire point. That paper's model solved far fewer problems on the first try than it did given more attempts: "we solve 70.2% of our problems with 100 samples per problem." Same model. Same difficulty. The only variable was how many shots you let it take. The authors called repeated sampling "a surprisingly effective strategy for producing working solutions to difficult prompts."
That is iteration, stated as arithmetic. More shots, more hits. It has been true of these models since before most people had heard of them.
There is one more finding I keep coming back to, because it is really about direction. Researchers kept asking whether a model can fix its own mistakes. The uncomfortable answer, from Google DeepMind in 2023: not on its own, not just by thinking harder about it. Left to introspect with no outside signal, models often made their answers worse. What fixed it was feedback, a real signal about whether the last attempt actually landed. You cannot correct toward nothing. You correct toward a target.
Hold onto that. It is the whole thing.
Forget the machines. This was always about you.
I did not write this much about model releases to talk about model releases. I wrote it because the machines just spent three years and unspeakable amounts of money rediscovering something a person gets for free, and most people are choosing not to take it.
A one-shot only works when there is no target.
That is the trick behind every one-shot demo. When the goal is "make something," anything you make counts, because nothing was aimed at. The moment you set a real direction, a specific place you are trying to reach, you introduce the possibility of missing. And anything that can be missed gets missed, usually the first time. That is not a defect in you. It is the geometry of aiming at something real.
So look at what actually has a direction in a life. A career. A company still standing in five years. A craft. A body. A name people trust. A piece of work you would put your signature on. Not one of these was ever going to arrive in a single clean motion, for the same reason no lab could bigger-model its way past the need to think twice: the target is real, and real targets get missed and re-approached. They get built the way everything worth anything gets built. Aim. Fire. See where it landed. Adjust. Fire again.
Your first shot missing is not failure. It is shot one. Do not read it as a verdict on you. Read it the way a good agent reads a failed test: as information. You now know something about where the target actually sits that you could not have known until you fired. That knowledge is the whole prize of the first shot. Missing is how you bought it.
The culture half-knows this already. The same crowd that says "one-shotted it" also adopted "vibe coding," which is just iteration wearing a friendlier face: nudge it, run it, feel what is off, nudge again. Even the meme quietly gave up on getting it right the first time.
Keep shooting
One of the most durable ideas in this field was written down back in 2019, before any of the current noise. Rich Sutton called it the bitter lesson: across seventy years of AI research, the approaches that win are the ones that lean on raw computation, and "the two methods that seem to scale arbitrarily in this way are search and learning." Search is trying many things and keeping what works. Learning is getting better from feedback. That is the machine's entire playbook. It is also yours. Try widely. Learn from what comes back. Neither of those is a single shot. Both of them are a person, with a direction, refusing to stop after the first miss.
So let the one-shot posts hype you up. They should. Watching something appear out of a single sentence is genuinely amazing. Just do not let the demo convince you the real thing is supposed to be that easy, because the real thing has a direction and the demo does not.
Pick your direction. Take the shot. When it misses, and it will, do not flinch and do not quit. Note where it landed, adjust, take the next one.
Keep shooting. With direction, iteration beats the one-shot every single time. The machines had to learn that the hard way, with all the money in the world. You already know it.
The one-shot was never real. You are.
This argument is where my name comes from. I call myself a LoopsHacker.
Sources
- OpenAI, "Learning to reason with LLMs" (o1), Sept 12, 2024: https://openai.com/index/learning-to-reason-with-llms/
- Anthropic, "Building Effective Agents," Dec 19, 2024: https://www.anthropic.com/research/building-effective-agents
- Anthropic, "Claude 3.7 Sonnet and Claude Code," Feb 24, 2025: https://www.anthropic.com/news/claude-3-7-sonnet
- Chen et al., "Evaluating Large Language Models Trained on Code" (pass@k), 2021: https://arxiv.org/abs/2107.03374
- Huang et al. (Google DeepMind), "Large Language Models Cannot Self-Correct Reasoning Yet," 2023: https://arxiv.org/abs/2310.01798
- METR, "Measuring AI Ability to Complete Long Tasks," 2025: https://arxiv.org/abs/2503.14499
- AutoGPT: https://en.wikipedia.org/wiki/AutoGPT
- Richard Sutton, "The Bitter Lesson," 2019: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Audio transcript
You are listening to "Keep Shooting," a note by Jain Yagi.
There is a fantasy hidden inside the phrase one shot.
You aim once. The answer arrives complete. The product works. The plan survives contact with reality. No miss, no correction, no evidence that the first idea was only a first idea.
For a while, that fantasy shaped how people talked about intelligent systems. Write the perfect prompt, press return, and expect the machine to land on the target immediately. When the result failed, the prompt was blamed or the model was dismissed.
But the systems that became genuinely useful did not become perfect at one shot. They learned to take more shots inside the process. Reasoning models spend time exploring before they commit. Effective agents plan, act, inspect the result, and revise. A run that looks autonomous from the outside is often thousands of small attempts, each redirected by what the previous one revealed.
This is not a special law of machines. It is the ordinary law of aiming.
The moment a task has a target, it becomes possible to miss. On the first attempt, you usually do. The miss tells you where the target is relative to the shot. That information is not failure. It is direction.
Benchmarks expose this clearly. A system may fail a problem on one attempt and solve it when allowed several independent attempts. The chance of at least one success rises because variation creates more paths to the answer. But repeated guessing is not enough for harder work. The powerful loop includes external feedback. Run the test. Read the error. Observe the user. Compare the result with the requirement. Then let reality aim the next shot.
Early autonomous agents often failed because they were given a large goal and expected to wander toward it with too little grounding. The agents that work better are less romantic. They operate in constrained loops. They use tools. They check state. They stop when evidence says stop.
I recognize the same pattern in my own work. A product is not one launch. A positioning sentence is not one draft. A design is not one act of taste. Each is a direction followed by a sequence of attempts. The advantage is not getting every first shot right. It is making the correction loop fast enough, honest enough, and cheap enough to continue.
Speed alone can create a machine gun pointed the wrong way. Persistence alone can repeat the same mistake forever. The loop needs a target and a signal. What are we trying to change? What did reality return? What should the next attempt do differently?
Intelligence helps choose the shot. Feedback gives the shot meaning. Iteration turns both into progress.
So keep shooting does not mean spray effort into the dark. It means choose a direction, take the shot, study the miss, and aim again. The target becomes clearer because you moved.
Nothing important gets one shot. That is not the tragedy of building. It is the method.


