한국어판: AI가 일을 더 싸게 만들수록, 병목은 어디로 이동할까
The most common story about AI is that it lets fewer people do more.
In many areas, that is already true enough to matter. Coding, drafting, analysis, search, and repetitive administrative work can all be accelerated by current systems.
But productivity gains are easy to misunderstand.
Making one stage of a process dramatically faster does not make the entire system equally fast.
A team can produce code faster without shipping software at the same rate. A government can automate form-filling without eliminating the need for verification. AI researchers can build more capable agents without solving the problem of evaluating and supervising them.
Across three very different conversations, the same analytical question keeps appearing:
When AI makes one kind of work cheaper, what becomes newly expensive?
That may be a more useful question than asking whether AI will simply “replace” a particular job.
When coding gets cheaper, choosing what to build becomes more valuable
In Y Combinator’s The State of Startups in 2026, the argument goes well beyond “AI makes programmers faster.”
YC says it is seeing more solo founders, smaller teams attempting larger problems, and a shift from pure software toward robotics, manufacturing, defense, semiconductors, photonics, power, and compute infrastructure.
The claim is not that AI builds factories or robots by itself.
It is that AI can reduce part of the software burden surrounding physical businesses. Internal tools, control software, prototypes, analysis, and workflow automation can require fewer people than before, making some ambitious projects accessible to smaller teams.
YC’s most interesting conclusion is also its most paradoxical:
As building becomes easier, knowing what to build becomes more important.
Knowing which customer problem is economically real, which request is merely noise, which market constraint will become painful later, and which idea should never be built in the first place is not automatically produced by better code generation.
That observation should not be treated as a universal fact. YC sees a highly selected population of technology companies built for rapid growth. Its experience is useful, but it is not representative of the average business.
Still, the underlying question can be tested against broader evidence.
Research shows a gap between producing code and shipping software
Evidence on AI coding productivity is not uniformly positive, but some field studies show substantial gains.
A 2026 Management Science paper pooled three randomized field experiments at Microsoft, Accenture, and another Fortune 100 company, covering 4,867 software developers. Across the combined sample, developers with access to an AI coding assistant completed about 26 percent more tasks on average.
That is meaningful evidence that AI can accelerate software work in at least some environments.
But producing more code is not the same as producing more finished software.
The September 2026 revision of NBER Working Paper 35275 by Mert Demirer, Leon Musolff, and Liyuan Yang analyzes AI usage and activity data from more than 500,000 GitHub developers.
The researchers found large increases in coding activity as tools progressed from autocomplete to interactive and autonomous agents. But those gains weakened sharply as they moved down the production chain.
For autonomous coding agents, the reported cumulative increase was 240 percent for commits, 80 percent for projects, and 30 percent for actual releases.
The pattern matters more than the headline number.
If code generation speeds up faster than review, testing, architecture, integration, and decision-making, those complementary activities become the constraint.
A system is not governed by its fastest step.
It is governed by how its dependent steps fit together.
The strongest counterargument: AI does not always make coding cheaper
There is a serious problem with any sweeping “AI makes software cheaper” thesis.
Sometimes it does not.
METR ran a randomized controlled trial in 2025 with 16 experienced open-source developers completing 246 real tasks in mature repositories they already knew well. With the early-2025 AI tools used in that setting, developers took 19 percent longer when AI was allowed.
Even more strikingly, they believed AI had made them faster.
That study does not prove that AI is bad for software development. METR explicitly warns against generalizing its findings to all developers, codebases, or future tools.
By early 2026, METR tried to repeat the experiment with newer systems, but strong selection effects made the results difficult to interpret. More developers were unwilling to participate if they had to work without AI, which biased the comparison.
The more defensible conclusion is narrower:
AI productivity is jagged.
It depends on the task, the user, the codebase, the generation of the tool, and the metric used to define success.
That uncertainty strengthens the bottleneck argument rather than weakening it. If gains are uneven, organizations need to know exactly where acceleration is occurring and where it is not.
Government bureaucracy shows the same cost-shifting problem
The idea becomes clearer outside software.
In her 2026 conversation with Tyler Cowen, journalist Annie Lowrey describes the time citizens spend navigating government systems as a “Time Tax.”
Tax filing, benefit applications, eligibility checks, waiting on hold, and repeatedly supplying the same information may not involve a direct payment. But they still consume time, and that time has an opportunity cost.
Lowrey’s argument is not simply that all bureaucracy should disappear.
Cowen presses the strongest objection: some friction has a purpose. Verification can reduce fraud. Screening can help allocate scarce services. Administrative difficulty can sometimes function as a form of rationing.
That changes the question.
It is no longer, “Should friction exist?”
It becomes:
Who should bear the cost of necessary verification?
If citizens repeatedly prove information the state already possesses, the cost is paid in citizen time. If systems pre-fill information, simplify front-end applications, automate repeated inputs, or rely more heavily on back-end auditing, part of that cost moves into the administrative system.
Again, the cost does not simply vanish.
Its location changes.
More capable AI makes evaluation and supervision more valuable
Noam Brown’s 2026 conversation with Dwarkesh Patel brings the same issue into AI research itself.
Brown discusses multi-agent systems, recent mathematical progress, and the possibility that AI could automate parts of AI research. Unlike human researchers, AI agents can be duplicated, run in parallel, and assigned different approaches to the same problem.
If that works at scale, research throughput could increase substantially.
But stronger capability creates another constraint:
How do we know the models are actually behaving as intended?
Brown emphasizes the gap between evaluation environments and real-world deployment. A system that performs well under a test may behave differently when given new capabilities, tools, or incentives outside that test.
Chain-of-thought monitoring is also not a complete solution. If visible reasoning is directly punished during training, a model might learn to make problematic reasoning less observable rather than eliminate the underlying objective.
Capability can therefore improve faster than oversight.
When that happens, evaluation becomes the bottleneck.
This is a lens, not a universal law
The three cases are not the same phenomenon.
A startup’s software pipeline, a government benefit system, and frontier AI evaluation operate under different incentives, institutions, and technical constraints.
Turning them into a grand rule—“every efficiency gain always creates a new cost”—would be too strong.
Some technologies really do eliminate old constraints almost completely. Some forms of automation reduce total costs without creating a comparably serious new problem.
The useful idea here is not a law.
It is a way to analyze change.
Whenever a technology produces a dramatic productivity gain, ask five questions:
- Which exact stage became faster?
- Did the final outcome improve at the same rate?
- If not, where did the gain attenuate?
- Which complementary human or organizational task became relatively more important?
- If that bottleneck is automated next, where does the constraint move after that?
Those questions are especially valuable with AI because the technology is changing multiple stages of work at different speeds.
The better question for the AI era
Organizations usually ask, “How much of this work can we automate?”
That is necessary, but incomplete.
If coding time falls by half while product judgment remains unchanged, product judgment becomes relatively more important.
If form-filling becomes automatic while eligibility rules remain difficult, verification remains the constraint.
If AI research accelerates while evaluation methods lag behind, oversight becomes the limiting factor.
A better question may be:
“If AI makes this part of the system dramatically cheaper, what becomes the next expensive thing?”
The value of AI is not just in producing more output.
It is in how quickly an organization can identify and redesign the next bottleneck created by the new speed.
Primary Sources
- Y Combinator — The State of Startups in 2026
- Annie Lowrey — Annie Lowrey on the Time Tax, Fraud, and Rationing by Paperwork
- Noam Brown — Agent swarms, alignment, & recursive self-improvement
- Cui et al. — The Effects of Generative AI on High-Skilled Work
- Demirer, Musolff & Yang — Writing Code vs. Shipping Code
- METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- METR — We are Changing our Developer Productivity Experiment Design
How the evidence is used
This article combines three source conversations with separate empirical research rather than treating them as a single body of evidence.
YC, Lowrey, and Brown are presented as attributed viewpoints. Developer-productivity findings are treated separately as empirical evidence. The argument that AI can shift bottlenecks rather than simply remove them is an interpretation that connects those sources.