An annotated reading guide · compiled 6 October 2026 arXiv:2604.15097 · cs.SE
Skill → Gene

Reading notes on From Procedural Skills to Strategy Genes — how LLM agents should encode experience so it actually controls behavior.

arXiv:2604.15097
Wang · Ren · Zhang
Submitted 16 Apr 2026

Findings · annotated

Every number in the paper, and what it is doing there

The Skill Probe, the Gene Probe, the Evolution Probe, and the two CritPt evolution runs: each table reproduced with its reading, its caveat, and its source.

Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package. Below, every table behind that sentence. Where the paper reports a sub-experiment with its own baseline, we say so rather than quietly mixing runs — one such case is the failure-ordering table near the end.

Skill Probe · 1,440 trialsDocumentation is not control

The opening move of the paper is almost rude: take the full Skill package, the artifact teams genuinely write and maintain, and check whether it helps. It does not. Skill lands a point under the no-guidance baseline on the average, because what it gives Flash it takes back from Pro with interest.

Table 2. Same experience source, three packagings. Δ in pp vs no guidance.
ConditionProFlashAvg.Δ
Gene59.9%48.2%54.0%+3.0
No guidance60.1%41.8%51.0%0.0
Skill50.7%49.0%49.9%−1.1

Source: Table 1 of arXiv:2604.15097. Avg. = arithmetic mean of the Pro and Flash pass rates.

Where the usable signal hides

Decomposing the Skill document section by section locates the problem precisely. The control value is not spread through the document; it is concentrated in one procedural slice. Skill-Workflow, alone, beats the baseline. Skill-Overview — the descriptive framing every documentation instinct tells you to write first — is the single most harmful component tested, four and a half points under baseline. The rest cluster near zero.

FIGURE 2. Skill decomposed. Sections are scored in isolation against the shared baseline; the full Skill package is shown for reference. Only the workflow slice clearly helps, and even it stays below a single distilled gene. Values as reported in the paper’s appendix (Table 12).

Brevity is not the whole story

A skeptic’s first move: maybe the Skill package loses simply because it is long. The paper tests this by truncating Skill fragments to roughly the gene’s 230-token budget. The fragments improve substantially — packaging overhead is real — but the gene still ends the comparison on top. Whatever the gene is doing, it is not merely being short.

Table 3. Budget-matched: Skill fragments cut to ≈230 tokens vs the Gene.
ConditionProFlashAvg.Δ
Gene59.9%48.2%54.0%+3.0
Skill — Pitfalls, short54.8%49.2%52.0%+1.0
Skill — Workflow, short54.6%48.3%51.5%+0.5
No guidance60.1%41.8%51.0%0.0

Source: Table 13 / §C.2 of arXiv:2604.15097. Fragments truncated to approximately the Gene’s prompt budget.

Gene Probe · 1,890 trialsAnatomy of the effect

The Gene Probe opens the gene up. Three questions: what does the gain emerge from, does it survive perturbation, and can it be topped up with extra material?

The gain arrives with strategy, not with tokens

Building the gene up piece by piece produces a result that should kill the “it’s just a short prompt” reading for good. Keywords alone: +2.5. Keywords plus a summary — strictly more content, strictly more tokens: back to zero. Full gene with the strategy layer: +3.0. The curve is not a ramp; it jumps exactly when the representation starts organizing experience into a control interface.

Table 4. Constructing the Gene from no guidance upward.
ConditionProFlashAvg.Δ
Gene (keywords + summary + strategy)59.9%48.2%54.0%+3.0
Gene (keywords only)57.9%49.1%53.5%+2.5
No guidance60.1%41.8%51.0%0.0
Gene (keywords + summary)51.3%50.6%51.0%0.0

Source: Table 2 of arXiv:2604.15097.

Robust to structure, sensitive to meaning

Mutating the gene separates two failure modes cleanly. Structural violence (inverting the priority order, over-constraining the guidance) barely dents it; one over-constrained variant actually scores above the clean gene. Semantic corruption — swapping in a wrong algorithm or a wrong domain — collapses it. And one mutation outperforms everything: the “stale paradigm” gene, carrying an outdated method that still frames the problem correctly, edges out the clean gene. The representation tolerates almost anything except losing touch with the task.

Table 5. Perturbed genes, sorted by average pass rate.
VariantProFlashAvg.
Stale paradigm — outdated method, right framing59.9%53.4%56.6%
Overconstrained57.2%54.5%55.9%
Clean Gene59.9%48.2%54.0%
Inverted priority54.8%50.8%52.8%
Wrong domain52.0%46.7%49.4%
Wrong algorithm49.6%47.9%48.8%

Source: Table 14 / §C.3 of arXiv:2604.15097. The paper reports no Δ column for these variants; none are invented here.

Documentation does not top up a gene — it dilutes it

If the gene were merely an incomplete document, reattaching the removed material should help. It does the opposite. API notes drag the gene half way back to baseline; examples cost another half point. The gene’s advantage is representational: once a compact control object is bloated back toward documentation, the extra text competes with the control signal it was supposed to carry.

Table 6. Reattaching documentation to the Gene.
ConditionProFlashAvg.Δ
Gene59.9%48.2%54.0%+3.0
Gene + examples57.8%46.1%52.0%+1.0
Gene + API notes51.8%51.2%51.5%+0.5
No guidance60.1%41.8%51.0%0.0
Skill50.7%49.0%49.9%−1.1

Source: Table 3 of arXiv:2604.15097.

Reuse has a scope boundary

The most practically inconvenient table in the paper is about composition. Naive bag-of-genes behavior (retrieve the k most related strategies and inject all of them) is not just unhelpful, it is the worst condition measured: two nominally complementary genes together fall 6.1 points below baseline, below even two conflicting genes. The reading we find most defensible: each additional partially-relevant control object blurs which lesson has authority over the current task. Selection is part of the reasoning problem; similarity search does not discharge it.

Table 7. Composing multiple Genes in one prompt.
ConditionProFlashAvg.Δ
Single Gene59.9%48.2%54.0%+3.0
Two conflicting Genes57.1%49.4%53.2%+2.2
No guidance60.1%41.8%51.0%0.0
Three complementary Genes54.5%46.2%50.4%−0.6
Two complementary Genes45.5%44.3%44.9%−6.1

Source: Table 4 of arXiv:2604.15097. The paper reads this as a scope boundary for high-difficulty scientific tasks, not as proof that genes are inherently non-compositional.

Evolution Probe · 1,260 trialsWhat accumulates well

The last probe treats the gene not as a one-shot prompt but as a carrier that experience attaches to over time. Three results define the carrier’s job description.

First, the carrier matters. The same failure history attached to a gene lands at 52.0%; the same material attached to a Skill package or freeform text lands under baseline. Accumulated experience is not carrier-neutral. Second, structure is not cosmetic: flatten the gene’s content into prose and most of its advantage evaporates (54.0% → 50.5%), even though nothing was removed. Third — and this is the line this guide would underline twice — attaching failure history naively still dilutes even the good carrier. What works is distillation: failure information compressed into standalone compact warnings outscores every mixed bundle the paper tries, including ordering failures before or after strategy.

Table 8. Encoding failure within the Gene — distilled warnings win.
ConditionProFlashAvg.Δ
Failure warnings only56.8%52.0%54.4%+4.6
Strategy only56.9%47.7%52.3%+2.5
Strategy first58.4%45.2%51.8%+2.0
Failure first56.3%44.7%50.5%+0.7
No guidance57.9%41.8%49.8%0.0

Source: Table 7 of arXiv:2604.15097. Mind the scope: this sub-experiment’s own baseline is 49.8%, not the 51.0% of the main comparison; Δ here is measured within this run.

Table 9. The same failure history, three carriers.
ConditionProFlashAvg.Δ
Gene59.9%48.2%54.0%+3.0
Gene + failure history55.3%48.6%52.0%+1.0
No guidance60.1%41.8%51.0%0.0
Freeform text + failure history55.7%43.5%49.6%−1.4
Skill + failure history53.8%41.8%47.8%−3.2

Source: Table 5 of arXiv:2604.15097. Flattened-prose comparison: Gene structured 54.0% vs Gene as prose 50.5% (Table 6 of the paper).

Test-time evolution · CritPtTwo days of evolution, weights untouched

Beyond the controlled probes, the paper reports two evolutionary runs on CritPt, a frontier-physics research benchmark. An evolutionary agent runs on OpenClaw as host runtime with Evolver, the evolution engine maintained by EvoMap, roughly two days per version, the base model fixed throughout. The first run (2026-02-16) is memory-grounded: it consolidates failures into reusable repair loops like gene_gep_repair_from_errors — structured diagnosis, blast-radius estimation, smallest reversible patch, validation, solidification. The second (2026-03-26) goes exploration-first, drawing on arXiv-derived and topic-prior genes, and banks procedural solution patterns — its most-selected high-value gene packages a Hamiltonian inverse-design procedure, and the recurring trio it forms with two companion genes appears across twelve tasks of the run.

Table 10. Gene-evolved systems vs their paired base models on CritPt.
Paired base modelBaseGene-evolvedGain
Gemini 3.1 Pro Preview (run 2026-03-26)17.7%27.14%+9.44 pp
Gemini 3 Pro Preview (run 2026-02-16)9.1%18.57%+9.47 pp

Endpoint figures as reported in the abstract of arXiv:2604.15097; the paper’s Appendix D documents the two runs, and the second run’s answers are published in a public repository it references (EvoMap/critpt-openclaw-reproducible-70). Gains shown in percentage points are the arithmetic difference of the two reported endpoints.

The mechanism, not the magnitude, is the transferable part. A stored gene can be inspected, diffed, versioned, tested against a sibling, and rolled back: governance that improvements buried in model weights do not admit. Two agents on one checkpoint can diverge into a pile of noisy summaries versus twenty validated strategy objects. Identical weights; different deployed systems.

MisreadingsFour readings this guide tries to prevent

Each of these circulates in discussion of the paper. Each is corrected by a specific number rather than an opinion, which is why they belong on the findings page.

M1“A Strategy Gene is just a shorter prompt.”

The budget-matched comparison answers this directly: Skill fragments cut to the gene’s ≈230-token budget improve to at most +1.0 pp, while the gene holds +3.0 pp — and flattening the gene’s own content into prose of identical content drops it from 54.0% to 50.5%. Length alone neither explains nor replaces the effect.

M2“If one gene helps, several should help more.”

The composition table says otherwise: two nominally complementary genes together score 44.9%, the worst condition measured anywhere in the paper — 6.1 pp under baseline and below the single-gene 54.0%. The paper’s framing is a scope boundary for specialized scientific tasks, but the direction is consistent: selection first, accumulation second.

M3“The stale-paradigm result means outdated methods are fine.”

The stale-paradigm gene (56.6%) carried an outdated method that still framed the problem correctly. The mutations that broke that framing collapsed: wrong domain 49.4%, wrong algorithm 48.8%. The finding is about problem framing surviving technical age, not a license to ship deprecated techniques.

M4“The CritPt gains will transfer to any agent.”

The abstract reports 9.1%→18.57% and 17.7%→27.14% on CritPt, a frontier-physics benchmark, in two specific paired setups, on a report the authors themselves label a beta technical report. The paper’s claim is that genes can support iterative improvement — not that double-digit lifts arrive everywhere. The controlled probes, not the CritPt endpoints, are the general evidence.

RecapThe four findings, in the paper’s own order

  1. Documentation-oriented skills are misaligned with test-time control: the useful signal is sparse, concentrated in a narrow procedural slice, and expanding a compact object into fuller documentation can degrade the average.
  2. Strategy genes are control-oriented objects, not shorter prompts: the benefit emerges with the strategy layer, survives structural perturbation, beats budget-matched fragments, and is diluted by reattached documentation.
  3. Gene is the better substrate for accumulation: attached failure history works best in a structured, editable carrier, and even then only selectively.
  4. Failure belongs in compact warnings, and evolution can bank increasingly useful strategy objects. Distilled warnings beat mixed bundles; two CritPt runs show gene-evolved systems pulling far ahead of their paired base models.

Every table above is quoted from the paper. To see how the two representations differ field by field, continue to Skill vs Gene; for the protocol that makes genes evolvable objects, see the GEP page.