'Beating RL With Reflection: GEPA and Optimize Anything - Lakshya A. Agrawal, GEPA'
Reflective optimization in text space - using rich textual feedback and natural language prompt updates instead of only scalar reward gradients - can dramaticall...
By Sean WeldonBeating RL With Reflection: GEPA and Optimize Anything
Abstract
This synthesis examines reflective optimization in text space, an alternative to weight-update methods (pretraining, supervised fine-tuning, reinforcement learning) for adapting AI systems to new tasks. The methodology substitutes scalar reward gradients with natural-language reflection on complete execution traces, enabling large, semantically coherent updates to textual artifacts. The GEPA algorithm operationalizes this through an evolutionary loop with Pareto-based candidate selection, reportedly achieving twice the performance gain of GRPO at 25,000 rollouts using only three data points in a single reflection round. Generalized under an "Optimize Anything" formulation, the approach extends beyond prompts to agent harnesses, kernels, skills, and scheduling policies, with reported gains including a 4.25%→30.52% improvement on AMD NPU code generation, ~40% cloud cost reduction, and 90x cost efficiency versus Claude Opus at Databricks. Implications for sample-efficient system adaptation are discussed, alongside open questions regarding evaluation criteria learning and continual improvement.
1. Introduction
Contemporary approaches to improving AI system behavior overwhelmingly rely on weight modification: pretraining consumes trillions of tokens, supervised fine-tuning (SFT) requires tens of thousands of labeled examples, and reinforcement learning (RL) with verified rewards may consume hundreds of thousands of rollouts. This data and compute regime assumes resources that most engineering teams do not possess. The problem is exacerbated as agentic systems mature: rollouts involving tool invocation, retrieval, and multi-step reasoning are increasingly long-running and expensive, rendering online RL infeasible for many practical deployments.
A second inefficiency compounds the first. In GRPO-style RL with verified rewards, an entire rollout - chain of thought, tool calls, intermediate outputs, compiler and runtime error messages - is compressed into a single binary or scalar reward. The diagnostic information a human engineer would use to debug a failure is discarded before it reaches the learning algorithm.
The central thesis examined here is that reflective optimization in text space - the use of an LLM to inspect complete execution traces and emit large natural-language edits to textual artifacts, rather than small numerical gradient updates - recovers this discarded information and thereby achieves substantial gains in sample efficiency. This thesis extends beyond prompt engineering: any artifact expressible as text and evaluable by a scoring function (agent code, harnesses, skills, kernels, scheduling policies) becomes optimizable under this paradigm. This analysis covers the theoretical motivation, the mechanics of the GEPA algorithm and its Pareto selection strategy, the generalized "Optimize Anything" framework and its reported empirical results, and implications for continual learning and evaluation design.
2. Background and Related Work
Reflective text optimization occupies a position between two established traditions. From reinforcement learning it inherits the loop structure of sampling, evaluation, and update. From evolutionary computation it inherits population-based search with mutation and selection. The distinguishing move is the substitution of the update operator itself: a natural-language rewrite replaces a gradient step.
The efficiency argument rests on the granularity of the update. Gradient descent applies many small deltas, each carrying limited semantic content, over thousands of steps. A textual edit can encode a large, coherent behavioral change in a single operation - for instance, changing an instruction from "generate a one-line summary" to "generate a 10-line summary" produces an immediate, substantial shift in output behavior that thousands of gradient steps might struggle to induce comparably. This positions text-space edits as high-leverage but discrete and non-differentiable, motivating search-based rather than descent-based optimization strategies such as those implemented in GEPA.
3. Core Analysis
3.1 The GEPA Algorithm: RL in Text Space
GEPA (Japa) functions as a reflective prompt optimization technique structured as an evolutionary loop with Pareto-based candidate selection. It resembles RL in that it receives a score for a candidate's performance, but unlike GRPO, it also receives rich textual feedback describing what succeeded and what failed in a given rollout. This reflection step can incorporate intermediate outputs and even additional tool calls, such as retrieval from a knowledge base, before producing an updated candidate.
The reported efficiency gains are substantial. In one round of reflection using only three data points, GEPA achieved twice the performance gain that GRPO obtained after 25,000 rollouts; continuing the optimization process doubled this gap again. Notably, Qwen 3 8B was reported to optimize itself with "no external expert teacher involved whatsoever," suggesting the mechanism does not require a stronger supervising model. Beyond raw score improvements, GEPA was observed to discover detailed problem specifications - purpose, context, and key lessons - rather than superficial prompt hacks, as evidenced by the AMD NPU XDNA2 case, where optimization pushed agent performance from 4.25% to 30.52% (a sevenfold improvement) purely through prompt refinement, including the discovery of domain-specific pitfalls such as "avoid including ADF.h."
3.2 Pareto Pool Selection Versus Loop-Based Search
A critical mechanistic finding concerns the candidate selection strategy. Simple LLM-in-a-loop search, which iteratively refines a single best candidate, was found to become trapped in local optima, repeatedly failing to improve from a plateaued node. GEPA's Pareto pool approach instead retains every candidate that wins on even one training example, rather than discarding all but the top scorer. Across four benchmarks, more than half of GEPA's total gains were attributable to this Pareto-based selection mechanism - approximately twice the gains achieved by loop-based optimization alone. This produced measured improvements of over 10% on question-answering, instruction-following, claim verification, and mathematical reasoning benchmarks, notably including domains where frontier labs have already invested heavily in optimization.
3.3 Generalization: Optimize Anything
The "Optimize Anything" framework generalizes the reflective optimization mechanism beyond prompts to any text-expressible, scorable artifact. The API requires three components: a set of problems, an evaluator or fitness function returning a score, and optional domain-specific "actionable side information" (compiler errors, profiler data, documentation, expert feedback). Three operating modes are defined: single-problem optimization, multitask search across related problems with information transfer, and a skill-building mode for generalization to novel queries at deployment time.
Reported results span diverse domains. RKGI accuracy for Gemini flash improved from 32.5% to 89.5% across 16 reflection rounds, with the optimizer discovering a six-step agent architecture involving rule hypothesis induction, code synthesis, execution and tracing, and debugging. Math500 accuracy for GPT-4.1 nano improved by 20% through a discovered two-step agent architecture. GSkill, a skill-learning feature, improved a MiniSU/GPT5 mini agent's repository issue resolution accuracy from 24% to 93%; transferring the learned skills to Claude Sonnet 4.5 achieved 100% accuracy with approximately 50% reduction in execution time and token usage. Cloud scheduling policy optimization reduced costs by approximately 40% relative to expert-designed heuristics, and OCR error rates for leading multimodal vision-language models were reduced by approximately 35% in externally validated tests.
3.4 Industrial Adoption and Cross-Model Transfer
Databricks reported a 90x cost reduction by tuning GPT OSS 120B via GEPA to outperform Claude Opus, with the improvement delta over Opus exceeding that observed on comparably-tuned open models - a finding the source material connects to the claim that smarter models benefit more from precise instructions, since they exhibit stronger instruction-following capacity. GEPA is reported to be used in production at Dropbox and Shopify, and is referenced in OpenAI's public discussion of self-improving AI systems, with zero hard dependencies enabling integration with arbitrary frameworks or models.
4. Technical Insights
Several implementation considerations follow from the reported findings. First, the amount of information supplied to the reflection step appears to determine result quality: "actionable side information" such as compiler diagnostics or profiler output materially improves optimization outcomes, suggesting practitioners should surface as much domain-specific signal as possible rather than relying on scalar scores alone. Second, the Pareto retention strategy implies that discarding low-scoring-but-locally-useful candidates during optimization removes a substantial fraction of achievable gains; systems implementing similar reflective loops should consider maintaining diverse candidate pools rather than greedy best-of selection. Third, the reported sample efficiency (three data points versus 25,000 rollouts) suggests this class of method is best suited to settings where rollouts are expensive or data is scarce, though the source material does not report controlled comparisons at matched compute budgets across all benchmarks, leaving open questions about generality. Fourth, the extension to skill-learning (GSkill) and cross-model transfer indicates that optimized artifacts may be more portable than fine-tuned weights, a property with direct implications for deployment cost.
5. Discussion
The reported findings suggest a broader reconsideration of where optimization effort should be allocated when adapting AI systems. If textual artifacts - prompts, harness code, scheduling policies - can be optimized with orders of magnitude fewer samples than weight updates require, the practical calculus for many engineering teams shifts toward reflective methods as a first resort rather than a fallback. This is particularly salient given the claim that smarter underlying models benefit more, not less, from precise textual instruction, which counters an intuition that model scaling might eventually obviate prompt-level optimization.
Nonetheless, several gaps merit further scrutiny. The reported comparisons between GEPA and GRPO are drawn from specific benchmark settings, and the source material does not provide a systematic account of failure modes or task classes where reflective optimization underperforms weight-based methods. The proposal to learn evaluation criteria themselves - optimizing LLM-as-judge prompts from approximately 50 annotated production traces - introduces a data flywheel in which judge quality and agent quality co-evolve, but this raises unaddressed questions about evaluator drift and validation. The "Learning Fast and Slow" direction, co-optimizing model weights and prompt harnesses jointly, points toward a potential unification of weight-based and text-based optimization for continual learning, though this remains an early-stage proposal rather than a validated system.
6. Conclusion
This analysis has examined reflective optimization in text space as a sample-efficient alternative to weight-update methods for adapting AI systems. The GEPA algorithm, through evolutionary search with Pareto-based candidate selection, demonstrates reported gains exceeding GRPO baselines by significant margins using minimal data, while the "Optimize Anything" generalization extends the approach to code, policies, and skills across diverse reported domains including hardware kernel generation, cloud scheduling, and multimodal OCR. The practical takeaway is that any artifact expressible as text and paired with a scoring function constitutes a candidate for this optimization paradigm, warranting evaluation by practitioners facing data- or rollout-constrained adaptation problems before resorting to full weight-update pipelines.
Sources
- Beating RL With Reflection: GEPA and Optimize Anything - Lakshya A. Agrawal, GEPA - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.