Annotated guide · Technical report
From skill to gene: the experience an agent keeps should control, not document.
A close reading of arXiv:2604.15097 — 4,590 controlled trials on why bulky Skill files make worse test-time guidance than a ~230-token Strategy Gene.
This guide walks through the paper’s question, its three probes, and every headline number, always checked against the text of the paper itself. Nothing here is a substitute for reading it; the notes are meant to make the read faster.
The questionWhat should an agent do with yesterday’s experience?
Agent systems have settled into a habit. When a run goes well — or badly — the instinct is to keep it: save the trajectory, write a retrospective, fold the lessons into a SKILL.md, grow the library. The assumption underneath is comfortable: a more complete record of experience should make the next attempt better.
The paper behind this guide arXiv:2604.15097 tests that assumption instead of inheriting it, and finds it points the wrong way. The useful question is not how much experience an agent stores. It is what form that experience takes when it re-enters the model at inference time. Stored documentation is written to be read by people. What the model needs at test time is a control signal — small, targeted, and hard to misread.
The authors put it in the abstract better than any paraphrase could:
“These results suggest that the core problem in experience reuse is not how to supply more experience, but how to encode experience as a compact, control-oriented, evolution-ready object.”
— arXiv:2604.15097, Abstract (Wang, Ren & Zhang, 2026)
Everything this site documents (the three probes, the tables, the Gene Evolution Protocol) is an unpacking of that one sentence. If you read nothing else, read the abstract on arXiv; it is unusually honest about scope, and it carries the two CritPt results (9.1%→18.57% and 17.7%→27.14%) that the evolutionary-agent section later earns.
The shiftSame experience, two encodings
The paper’s central comparison takes one source of experience per scenario and recasts it two ways: as a full documentation-oriented Skill package (~2,500 tokens), and as a compact Strategy Gene (~230 tokens). The content overlaps; the packaging is what differs — and the packaging is what moves the score.
Skill · documentation package · ≈2,500 tokens
useful signal: sparse — one narrow procedural slice (Skill Probe, §4.1)
Strategy Gene · control object · ≈230 tokens
control: compact, behavior-targeted, failure-aware
Two details make this more than a compression trick. First, the gene is not a shortened skill; the paper says so outright. It is a different abstraction: organized by control logic (when it applies, what to do, what to avoid) rather than documentation logic (what explains the task to a human). Second, the benefit is not explained by token count alone: when the authors trimmed Skill fragments down to the gene’s budget, the fragments got better but still lost to the gene. We unpack that on the Skill vs Gene page.
MethodThree probes, 4,590 trials
The evidence is organized as three probes on Gene-Bench — the paper’s name for its 45-scenario scientific code-solving benchmark. Every scenario generates a Python program that is executed in a sandbox and scored by checkpoints, so partial progress counts and no single test suite dominates the average.
| Probe | Retained trials | Question it answers | Headline finding |
|---|---|---|---|
| Skill Probe | 1,440 | Why documentation packages fail as test-time control | Skill sits below the no-guidance baseline; only the Workflow section clearly helps |
| Gene Probe | 1,890 | Whether Gene is a real representation, not a short prompt | Gene is strongest overall, robust to structural perturbation, diluted by reattached docs |
| Evolution Probe | 1,260 | Which carrier supports accumulation over time | Structured genes carry failure history best; compact warnings beat appended logs |
Counts as reported in the paper (§4.1, §4.2, §4.3); they sum to the 4,590 retained trials stated in the abstract.
Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package. That single sentence is the spine of the paper; the rest of the findings page walks through each table that supports it, with every number traced back to the source.
Headline result+3.0 points from packaging alone
The cleanest comparison in the paper holds the experience source fixed and varies only the representation. The full Skill package, the thing our instincts tell us to write, lands below the baseline: it lifts the smaller Flash model but costs the stronger Pro model more than it gains. The gene, at roughly a tenth of the tokens, is the only condition that beats no guidance on the average.
Average pass rate, Pro and Flash combined · Source: Table 1, arXiv:2604.15097
Watch the Pro column, though. Pro alone actually drops slightly under the gene (60.1% → 59.9%) while Flash rises sharply (41.8% → 48.2%). The honest summary is not “genes make strong models better”; it is “the representation decides where guidance helps and where it hurts” — which is, if anything, a more useful result for anyone tuning an agent.
Beyond one shotGenes as an interface for evolution
The second half of the paper asks a harder question than one-shot control: if experience lives in discrete, structured objects, can a system improve across runs without touching the model weights? To test it the authors stand an evolutionary agent on OpenClaw as host runtime with Evolver as the evolution engine, and let two versions run on CritPt, a frontier-physics benchmark — roughly two days of evolution each.
The gains are the ones quoted in the abstract: a gene-evolved system lifted from 9.1% to 18.57% over its paired base model, and a later version from 17.7% to 27.14%. Just as interesting is what the two runs accumulated. The first distills failures into reusable repair loops — a gene that packages structured diagnosis, blast-radius estimation, the smallest reversible patch, validation, solidification. The second goes further and banks procedural solution patterns from run experience and arXiv-derived genes, then re-deploys them at scale: one Hamiltonian inverse-design gene alone was selected repeatedly across the run.
A useful mental model, borrowed from the community discussion rather than the paper: the rich skill is source code, the gene is the compiled artifact; compilation is what makes versioning, validation, and rollback possible. The protocol that formalizes this is the Gene Evolution Protocol.
FAQQuestions this guide actually gets
Q1What does “skill to gene” mean?
On this site, “skill to gene” refers to the representational shift studied in arXiv:2604.15097 — replacing documentation-oriented Skill packages with compact, control-oriented Strategy Gene objects as the way LLM agents carry reusable experience. It is not a term from biology or genetics, and it is not a product name.
Q2What is a Strategy Gene?
A Strategy Gene is a compact, structured, control-oriented representation of an agent’s reusable experience, distilled from richer sources such as procedural skills or past trajectories. In the paper’s operational setup it is roughly 230 tokens long and carries domain keywords, a one-sentence summary, a short strategy list, and failure-aware AVOID cues, with optional constraint and validation fields.
Q3What is the main result of arXiv:2604.15097?
Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package. On the CritPt benchmark, gene-evolved systems improved over their paired base models from 9.1% to 18.57% and from 17.7% to 27.14%.
Q4Is a Strategy Gene just a shorter prompt?
No. When Skill fragments were trimmed to match the Gene’s token budget, they improved but still stayed below the Gene’s 54.0% average pass rate. The keywords-only Gene variant already scored 53.5%, while flattening the structured Gene into prose of the same content dropped it to 50.5% — the structure itself, not just brevity, carries the effect.
Q5What is the Gene Evolution Protocol (GEP)?
The Gene Evolution Protocol is the protocol layer introduced in the paper to canonicalize genes into structured, evolvable objects. It organizes experience into three object types — genes (control units), capsules (validated execution paths), and events (immutable evolution logs) — and updates them through a six-stage loop: Scan, Signal, Intent, Mutate, Validate, Solidify.
Q6Which models and tasks were used in the experiments?
All retained experiments used two fixed models, Gemini 3.1 Pro Preview and Gemini 3.1 Flash Lite Preview, on 45 scientific code-solving scenarios spanning bioinformatics, neuroscience, chemistry, seismology, climate science, signal processing, network analysis, finance, robotics, and quantum computing. Generated programs were executed in a sandbox and scored by checkpoint-based pass rate.
Q7Where can I read the paper?
The paper is available on arXiv at https://arxiv.org/abs/2604.15097, with a PDF at https://arxiv.org/pdf/2604.15097, an experimental HTML version at https://arxiv.org/html/2604.15097v1, and the DOI 10.48550/arXiv.2604.15097. Ready-to-use BibTeX is on our cite page.
Suggested reading path
- Findings — every table, annotated
- Skill vs Gene — the two representations, field by field
- GEP Protocol — genes, capsules, events
- Gene-Bench — the 45 scenarios and the metric
- Cite — BibTeX and identifiers
- Commentary — discussion elsewhere