An annotated reading guide · compiled 6 October 2026 arXiv:2604.15097 · cs.SE
Skill → Gene

Reading notes on From Procedural Skills to Strategy Genes — how LLM agents should encode experience so it actually controls behavior.

arXiv:2604.15097
Wang · Ren · Zhang
Submitted 16 Apr 2026

Concepts · definitions first

Skill vs Gene: a documentation artifact and a control object

A procedural Skill is a documentation-oriented experience representation: a human-readable package that organizes prior problem-solving knowledge into overview, workflow, API notes, examples, error handling, and pitfalls. It is built to be read, taught, reviewed, and archived.

A Strategy Gene is a control-oriented experience representation: a compact object distilled from that same experience, carrying task-matching keywords, a one-sentence summary, a short strategy list, and failure-aware AVOID cues, built to change what the model does on the next run under a tight token budget.

Those two sentences carry the whole page. In arXiv:2604.15097, the same underlying experience is cast into both forms and compared across 4,590 controlled trials: the gene ends at a 54.0% average pass rate, the full Skill at 49.9%, no guidance at 51.0%. The argument of this page is that the gap is neither noise nor brevity; it follows from what each encoding is for.

Side by sideEvery axis on which they differ

Table 1. The two representations, as characterized in the paper (§3.2, §3.3, Appendix A).
AxisProcedural SkillStrategy Gene
OrientationDocumentation-oriented — written for human readingControl-oriented — written for model-facing inference
GoalDocumentary completeness; teach and archive the processSignal density under a constrained token budget
OrganizationDocumentation logic: overview → workflow → reference materialControl logic: when it applies → what to do → what to avoid → how to check
Failure knowledgeRecorded as history — error logs, pitfalls sections, worked failuresCompressed into explicit AVOID cues attached to the strategy
Structure’s roleFormatting serves readabilityEditable schema is part of the effect — flattened to prose, the gain collapses (54.0% → 50.5%)
AccumulationGrows by appending; history dilutes control (Skill + failure: 47.8%)Evolves by selective revision; validated warnings attach cleanly (Gene + failure: 52.0%)
Natural consumerA developer six months laterA model in the next run

Percentages from the corresponding tables of arXiv:2604.15097, annotated on the findings page.

AnatomyWhat is actually inside each one

Concretely, for the paper’s running example (scenario S012_uv_spectroscopy — detect and measure peaks in UV-Vis spectra): the Skill side is the full package a maintainer would recognize, seven sections deep. The gene is eight lines. Both encode the same hard-won lesson about unit conversion; only one of them surfaces it as an instruction the model cannot miss.

The recurring failure the gene guards against is easy to underestimate. scipy.signal.find_peaks expects min_distance in sample-index units; a model that passes a wavelength value through unconverted produces plausible-looking output with wrong peak counts. Separately, reporting FWHM requires converting peak_widths output back to wavelength units first. Neither error breaks the program loudly. Both quietly cost checkpoints — which is exactly the kind of failure a compact AVOID cue is for.

SKILL.md · excerpt, ≈2,500 tokens in full

strategy-gene · complete, ≈230 tokens

<strategy-gene> Domain keywords: uv-vis, peak detection, FWHM, unit conversion Summary: Detect peaks and compute wavelength-domain peak properties correctly Strategy: 1. Detect peaks with prominence-based criteria 2. Convert min_distance into sample-index units before peak detection 3. AVOID: Report FWHM only after converting peak_widths outputs back to wavelength units </strategy-gene>
FIGURE 1. The same lesson in both encodings. The Skill buries the unit-conversion constraint inside API notes and pitfalls sections; the gene promotes it to a numbered strategy step and an AVOID cue. Gene content reproduced from §3.3.2 of the paper.

The gene schema, field by field

Under the Gene Evolution Protocol, the gene is serialized as a structured object. The paper (Appendix A.3) lists the fields: type, schema_version, id, signals_match — the keywords that decide when the gene applies — then summary, strategy, optional constraints, optional validation hooks, and an asset_id for lineage. Each field earns its place in the controlled trials: keywords alone carry +2.5, but the jump to the strongest average comes when the strategy layer completes the object.

What the schema buys is operability. Genes with stable boundaries can be matched against a task, replaced by a revised version, composed deliberately rather than heaped, and validated before solidification. A prose blob can be none of those things — it can only be re-read, or re-generated. That is the precise sense in which a gene is not a shortened skill. It is the smallest unit of experience a system can operate on rather than merely quote.

Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package. The anatomy above is the paper’s explanation of why the small object wins.

For the protocol layer that manages genes over time — capsules, events, the six-stage loop — continue to the GEP page. For every table behind the percentages on this page, see findings.