AI-Generated Code Is Already Competing With Human Code - Daksh Gupta, Greptile
Data from Graphite's review of over a million monthly pull requests shows that fully autonomous AI coding agents now produce a growing share of enterprise co...
By Sean WeldonAI-Generated Code Is Already Competing With Human Code
Abstract
Empirical telemetry from a code review platform processing more than one million pull requests (PRs) per month across thousands of companies indicates that fully autonomous artificial intelligence (AI) coding agents have transitioned from experimental curiosity to material contributor in enterprise codebases. Detection heuristics combining PR description footers and branch-name prefixes suggest that approximately 25% of monthly reviewed PRs are now fully or largely machine-authored, up from under 1% roughly a year earlier. Across three independent quality proxies - revert rates, severity-graded bug comments, and review iteration counts - agent-authored changes are statistically comparable to human-authored changes. However, keyword-level analysis of millions of review comments reveals agent-specific failure distributions, including elevated SQL injection incidence for one model family. These findings motivate a first-principles reformulation of code validation suited to PR volumes that manual review cannot absorb.
1. Introduction
The question of whether end-to-end coding agents can contribute to commercially viable codebases has, until recently, lacked large-scale empirical grounding. Public discourse has been dominated by anecdote - reports of agents opening 100 PRs per day, or developers adopting polyphasic sleep schedules to supervise agent cycles - rather than by measurement against production repositories with real customers.
This synthesis examines evidence drawn from Graphite, a code review platform used weekly by tens of thousands of engineers and reviewing in excess of one million PRs monthly. Three research questions structure the analysis: What proportion of enterprise PR volume is now authored by autonomous agents, and how is that proportion changing? Does agent-authored code differ measurably in quality from human-authored code? If quality is comparable but volume is not, what validation architecture can scale accordingly?
Key terminology is defined as follows. A fully autonomous coding agent is a system capable of completing an entire pull request - problem interpretation, multi-file modification, and submission - without intermediate human authorship. A pull request (PR) is the atomic unit of proposed change analyzed throughout. Revert rate denotes the fraction of merged PRs subsequently undone, used here as a post-hoc quality proxy. The central thesis is that agent-authored code is no longer a marginal phenomenon but a statistically comparable, qualitatively distinct contributor to enterprise codebases, necessitating new validation infrastructure.
2. Background and Related Work
The capability trajectory of AI-assisted software development segments into three phases. In 2022, the release of GPT-3.5 established tab-completion as the dominant interaction paradigm, instantiated commercially in Cursor and GitHub Copilot. In 2024, reliable multi-file editing emerged, pioneered by Cursor, extending model influence beyond single-buffer context. In 2025, fully autonomous agents capable of closing an entire PR became viable; December of the preceding year is identified as the watershed at which agents crossed into practical autonomy.
Notably, the observed growth in agent-authored PR share has been continuous rather than discontinuous. It does not correlate tightly with individual frontier model releases, suggesting diffusion is governed by organizational adoption dynamics and tooling maturity rather than discrete capability jumps. This distinction matters for forecasting: adoption curves driven by workflow integration tend to be smoother and more persistent than those driven by model announcements, implying that measurement based on vendor release calendars will systematically lag reality.
3. Core Analysis
3.1 Measuring Agent Authorship
Attribution is methodologically non-trivial. Querying the GitHub author field for Codex, Claude, or Cursor identifies fewer than 1% of PRs as agent-authored - an unreliable lower bound, since agents frequently commit under the credentials of the invoking human. Two supplementary signals materially improve recall: PR description footers (e.g., "co-authored by Claude") and branch-name prefixes (e.g., Codex-named branches). Combined, these signals indicate that approximately 25% of PRs reviewed monthly by Graphite are fully or largely AI-generated, a rise from under 1% over roughly one year. The trajectory's continuity, rather than its stepwise correlation with model releases, supports the interpretation that enterprise adoption is now structurally embedded rather than experimentally episodic.
3.2 Quality Comparison Across Proxies
Three independent proxies were used to assess relative quality. Revert rate, computed by exploiting GitHub's convention of naming revert branches revert-PR number-PR name, showed Codex at approximately 1 per 1,000 PRs, Devon at approximately 3.5 per 1,000, and humans at approximately 2.5 per 1,000. No significant correlation was found between PR size and revert rate for either humans or agents, suggesting that revert likelihood is not simply a function of change magnitude. Second, bug severity classification (P0/P1/P2) derived from Graphite's review comments showed that three of four examined agents produced fewer P0 (highest-severity) bugs than humans. Third, review iteration counts prior to merge were comparable across populations: Devon averaged 2.1 cycles, Codex 2.45, with human-authored PRs falling in a similar range. Taken together, these three proxies yield little statistical evidence that AI-generated PRs are of lower quality than human-generated ones - a finding that runs counter to the skepticism that end-to-end agents remain unsuitable for commercially viable codebases.
3.3 Divergent Failure Patterns
Although aggregate quality metrics are comparable, qualitative analysis of failure types reveals meaningful divergence. Scanning millions of Graphite review comments for specific defect terms (e.g., SQL injection, N+1 query, auth bypass) showed that Claude is approximately 1.5 times more likely than humans to produce SQL injection vulnerabilities, while Devon is approximately half as likely as humans to introduce authentication bypass issues. This indicates that model-specific failure signatures exist even where overall defect rates are similar, implying that quality assurance tuned to a single aggregate metric would miss systematic, model-specific risk concentrations.
3.4 Scale of Enterprise PR Volume
The practical urgency of these findings is reinforced by volume statistics. The median Graphite user submits 50 PRs per month (just over two per workday), the 90th percentile submits 500 PRs per month, and the 99th percentile submits thousands per month. The gap between median and P90 behavior is substantial, indicating a long tail of extremely high-throughput users - plausibly agent-driven - for whom existing manual QA and code review processes were not designed.
4. Technical Insights
Several implementation-relevant findings emerge. First, attribution heuristics matter: GitHub author-field detection alone undercounts agent authorship by more than an order of magnitude, and robust measurement requires combining metadata signals (PR footers, branch-naming conventions). Second, revert rate as a proxy benefits from GitHub's branch-naming convention but is coarse-grained and insensitive to PR size effects, limiting its diagnostic precision. Third, the ~4 comments per PR generated on average by Graphite's review process constitutes a corpus of millions of comments, enabling keyword-level failure-mode analysis that aggregate metrics (revert rate, severity) cannot surface. Fourth, a proposed three-question validation framework - does the change violate the user contract, does it increase future violation risk, does it fulfill the author's intent - offers a compact decision structure for automated merge gating. Fifth, agent-based validation itself (sandboxed execution, dependency installation, input mocking, browser-agent testing) is presented as a mechanism to close the loop without proportional human review. A key trade-off is that nearly 20% of PRs are currently merged without any human review or testing, raising the stakes for the reliability of automated validation substitutes.
5. Discussion
These findings reframe the central question from whether agents can produce acceptable code to how validation infrastructure should be redesigned given that they already do so at scale. The comparable aggregate quality metrics suggest that resistance to agent-authored code should not rest on a presumption of inferior craftsmanship. However, the divergent failure patterns - elevated SQL injection risk for one model, reduced auth-bypass risk for another - indicate that quality assurance must become model-aware rather than treating "AI-generated" as a monolithic category.
A significant knowledge gap concerns causality: it remains unclear whether these failure-pattern differences stem from training data composition, agent architecture, or task-selection effects (e.g., which PR types are routed to which agents). Additionally, the reliance on keyword-based scanning for failure-mode detection is inherently limited to named, anticipated defect categories, and may miss novel failure classes unique to agentic code generation. The scale mismatch between PR throughput (particularly at the P90/P99 tail) and manual review capacity situates this work within a broader industry trend toward automated, contract-based validation rather than line-by-line human inspection.
6. Conclusion
This analysis demonstrates that fully autonomous coding agents have moved from anecdotal novelty to a measurable, substantial share - approximately 25% - of enterprise pull request volume, with quality metrics statistically comparable to human-authored code despite distinct, model-specific failure signatures. The practical takeaway is that validation strategy must shift from authorship-based skepticism toward content-based, model-aware risk assessment, operationalized through frameworks such as the three-question contract-violation model. Future work should prioritize refining attribution methodology, expanding failure-mode taxonomies beyond keyword matching, and empirically testing automated merge-gating systems against the high-throughput tail of enterprise PR activity.
Sources
- AI-Generated Code Is Already Competing With Human Code - Daksh Gupta, Greptile - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.