Findings · annotated
Every number in the paper, and what it is doing there
The Skill Probe, the Gene Probe, the Evolution Probe, and the two CritPt evolution runs: each table reproduced with its reading, its caveat, and its source.
Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package. Below, every table behind that sentence. Where the paper reports a sub-experiment with its own baseline, we say so rather than quietly mixing runs — one such case is the failure-ordering table near the end.
Skill Probe · 1,440 trialsDocumentation is not control
The opening move of the paper is almost rude: take the full Skill package, the artifact teams genuinely write and maintain, and check whether it helps. It does not. Skill lands a point under the no-guidance baseline on the average, because what it gives Flash it takes back from Pro with interest.
| Condition | Pro | Flash | Avg. | Δ |
|---|---|---|---|---|
| Gene | 59.9% | 48.2% | 54.0% | +3.0 |
| No guidance | 60.1% | 41.8% | 51.0% | 0.0 |
| Skill | 50.7% | 49.0% | 49.9% | −1.1 |
Source: Table 1 of arXiv:2604.15097. Avg. = arithmetic mean of the Pro and Flash pass rates.
Where the usable signal hides
Decomposing the Skill document section by section locates the problem precisely. The control value is not spread through the document; it is concentrated in one procedural slice. Skill-Workflow, alone, beats the baseline. Skill-Overview — the descriptive framing every documentation instinct tells you to write first — is the single most harmful component tested, four and a half points under baseline. The rest cluster near zero.
│ reference mark = no-guidance baseline (51.0%) · avg. pass rate, axis 44–58% · Source: Table 12 / §C.1 of arXiv:2604.15097
Brevity is not the whole story
A skeptic’s first move: maybe the Skill package loses simply because it is long. The paper tests this by truncating Skill fragments to roughly the gene’s 230-token budget. The fragments improve substantially — packaging overhead is real — but the gene still ends the comparison on top. Whatever the gene is doing, it is not merely being short.
| Condition | Pro | Flash | Avg. | Δ |
|---|---|---|---|---|
| Gene | 59.9% | 48.2% | 54.0% | +3.0 |
| Skill — Pitfalls, short | 54.8% | 49.2% | 52.0% | +1.0 |
| Skill — Workflow, short | 54.6% | 48.3% | 51.5% | +0.5 |
| No guidance | 60.1% | 41.8% | 51.0% | 0.0 |
Source: Table 13 / §C.2 of arXiv:2604.15097. Fragments truncated to approximately the Gene’s prompt budget.
Gene Probe · 1,890 trialsAnatomy of the effect
The Gene Probe opens the gene up. Three questions: what does the gain emerge from, does it survive perturbation, and can it be topped up with extra material?
The gain arrives with strategy, not with tokens
Building the gene up piece by piece produces a result that should kill the “it’s just a short prompt” reading for good. Keywords alone: +2.5. Keywords plus a summary — strictly more content, strictly more tokens: back to zero. Full gene with the strategy layer: +3.0. The curve is not a ramp; it jumps exactly when the representation starts organizing experience into a control interface.
| Condition | Pro | Flash | Avg. | Δ |
|---|---|---|---|---|
| Gene (keywords + summary + strategy) | 59.9% | 48.2% | 54.0% | +3.0 |
| Gene (keywords only) | 57.9% | 49.1% | 53.5% | +2.5 |
| No guidance | 60.1% | 41.8% | 51.0% | 0.0 |
| Gene (keywords + summary) | 51.3% | 50.6% | 51.0% | 0.0 |
Source: Table 2 of arXiv:2604.15097.
Robust to structure, sensitive to meaning
Mutating the gene separates two failure modes cleanly. Structural violence (inverting the priority order, over-constraining the guidance) barely dents it; one over-constrained variant actually scores above the clean gene. Semantic corruption — swapping in a wrong algorithm or a wrong domain — collapses it. And one mutation outperforms everything: the “stale paradigm” gene, carrying an outdated method that still frames the problem correctly, edges out the clean gene. The representation tolerates almost anything except losing touch with the task.
| Variant | Pro | Flash | Avg. |
|---|---|---|---|
| Stale paradigm — outdated method, right framing | 59.9% | 53.4% | 56.6% |
| Overconstrained | 57.2% | 54.5% | 55.9% |
| Clean Gene | 59.9% | 48.2% | 54.0% |
| Inverted priority | 54.8% | 50.8% | 52.8% |
| Wrong domain | 52.0% | 46.7% | 49.4% |
| Wrong algorithm | 49.6% | 47.9% | 48.8% |
Source: Table 14 / §C.3 of arXiv:2604.15097. The paper reports no Δ column for these variants; none are invented here.
Documentation does not top up a gene — it dilutes it
If the gene were merely an incomplete document, reattaching the removed material should help. It does the opposite. API notes drag the gene half way back to baseline; examples cost another half point. The gene’s advantage is representational: once a compact control object is bloated back toward documentation, the extra text competes with the control signal it was supposed to carry.
| Condition | Pro | Flash | Avg. | Δ |
|---|---|---|---|---|
| Gene | 59.9% | 48.2% | 54.0% | +3.0 |
| Gene + examples | 57.8% | 46.1% | 52.0% | +1.0 |
| Gene + API notes | 51.8% | 51.2% | 51.5% | +0.5 |
| No guidance | 60.1% | 41.8% | 51.0% | 0.0 |
| Skill | 50.7% | 49.0% | 49.9% | −1.1 |
Source: Table 3 of arXiv:2604.15097.
Reuse has a scope boundary
The most practically inconvenient table in the paper is about composition. Naive bag-of-genes behavior (retrieve the k most related strategies and inject all of them) is not just unhelpful, it is the worst condition measured: two nominally complementary genes together fall 6.1 points below baseline, below even two conflicting genes. The reading we find most defensible: each additional partially-relevant control object blurs which lesson has authority over the current task. Selection is part of the reasoning problem; similarity search does not discharge it.
| Condition | Pro | Flash | Avg. | Δ |
|---|---|---|---|---|
| Single Gene | 59.9% | 48.2% | 54.0% | +3.0 |
| Two conflicting Genes | 57.1% | 49.4% | 53.2% | +2.2 |
| No guidance | 60.1% | 41.8% | 51.0% | 0.0 |
| Three complementary Genes | 54.5% | 46.2% | 50.4% | −0.6 |
| Two complementary Genes | 45.5% | 44.3% | 44.9% | −6.1 |
Source: Table 4 of arXiv:2604.15097. The paper reads this as a scope boundary for high-difficulty scientific tasks, not as proof that genes are inherently non-compositional.
Evolution Probe · 1,260 trialsWhat accumulates well
The last probe treats the gene not as a one-shot prompt but as a carrier that experience attaches to over time. Three results define the carrier’s job description.
First, the carrier matters. The same failure history attached to a gene lands at 52.0%; the same material attached to a Skill package or freeform text lands under baseline. Accumulated experience is not carrier-neutral. Second, structure is not cosmetic: flatten the gene’s content into prose and most of its advantage evaporates (54.0% → 50.5%), even though nothing was removed. Third — and this is the line this guide would underline twice — attaching failure history naively still dilutes even the good carrier. What works is distillation: failure information compressed into standalone compact warnings outscores every mixed bundle the paper tries, including ordering failures before or after strategy.
| Condition | Pro | Flash | Avg. | Δ |
|---|---|---|---|---|
| Failure warnings only | 56.8% | 52.0% | 54.4% | +4.6 |
| Strategy only | 56.9% | 47.7% | 52.3% | +2.5 |
| Strategy first | 58.4% | 45.2% | 51.8% | +2.0 |
| Failure first | 56.3% | 44.7% | 50.5% | +0.7 |
| No guidance | 57.9% | 41.8% | 49.8% | 0.0 |
Source: Table 7 of arXiv:2604.15097. Mind the scope: this sub-experiment’s own baseline is 49.8%, not the 51.0% of the main comparison; Δ here is measured within this run.
| Condition | Pro | Flash | Avg. | Δ |
|---|---|---|---|---|
| Gene | 59.9% | 48.2% | 54.0% | +3.0 |
| Gene + failure history | 55.3% | 48.6% | 52.0% | +1.0 |
| No guidance | 60.1% | 41.8% | 51.0% | 0.0 |
| Freeform text + failure history | 55.7% | 43.5% | 49.6% | −1.4 |
| Skill + failure history | 53.8% | 41.8% | 47.8% | −3.2 |
Source: Table 5 of arXiv:2604.15097. Flattened-prose comparison: Gene structured 54.0% vs Gene as prose 50.5% (Table 6 of the paper).
Test-time evolution · CritPtTwo days of evolution, weights untouched
Beyond the controlled probes, the paper reports two evolutionary runs on CritPt, a frontier-physics research benchmark. An evolutionary agent runs on OpenClaw as host runtime with Evolver, the evolution engine maintained by EvoMap, roughly two days per version, the base model fixed throughout. The first run (2026-02-16) is memory-grounded: it consolidates failures into reusable repair loops like gene_gep_repair_from_errors — structured diagnosis, blast-radius estimation, smallest reversible patch, validation, solidification. The second (2026-03-26) goes exploration-first, drawing on arXiv-derived and topic-prior genes, and banks procedural solution patterns — its most-selected high-value gene packages a Hamiltonian inverse-design procedure, and the recurring trio it forms with two companion genes appears across twelve tasks of the run.
| Paired base model | Base | Gene-evolved | Gain |
|---|---|---|---|
| Gemini 3.1 Pro Preview (run 2026-03-26) | 17.7% | 27.14% | +9.44 pp |
| Gemini 3 Pro Preview (run 2026-02-16) | 9.1% | 18.57% | +9.47 pp |
Endpoint figures as reported in the abstract of arXiv:2604.15097; the paper’s Appendix D documents the two runs, and the second run’s answers are published in a public repository it references (EvoMap/critpt-openclaw-reproducible-70). Gains shown in percentage points are the arithmetic difference of the two reported endpoints.
The mechanism, not the magnitude, is the transferable part. A stored gene can be inspected, diffed, versioned, tested against a sibling, and rolled back: governance that improvements buried in model weights do not admit. Two agents on one checkpoint can diverge into a pile of noisy summaries versus twenty validated strategy objects. Identical weights; different deployed systems.
MisreadingsFour readings this guide tries to prevent
Each of these circulates in discussion of the paper. Each is corrected by a specific number rather than an opinion, which is why they belong on the findings page.
M1“A Strategy Gene is just a shorter prompt.”
The budget-matched comparison answers this directly: Skill fragments cut to the gene’s ≈230-token budget improve to at most +1.0 pp, while the gene holds +3.0 pp — and flattening the gene’s own content into prose of identical content drops it from 54.0% to 50.5%. Length alone neither explains nor replaces the effect.
M2“If one gene helps, several should help more.”
The composition table says otherwise: two nominally complementary genes together score 44.9%, the worst condition measured anywhere in the paper — 6.1 pp under baseline and below the single-gene 54.0%. The paper’s framing is a scope boundary for specialized scientific tasks, but the direction is consistent: selection first, accumulation second.
M3“The stale-paradigm result means outdated methods are fine.”
The stale-paradigm gene (56.6%) carried an outdated method that still framed the problem correctly. The mutations that broke that framing collapsed: wrong domain 49.4%, wrong algorithm 48.8%. The finding is about problem framing surviving technical age, not a license to ship deprecated techniques.
M4“The CritPt gains will transfer to any agent.”
The abstract reports 9.1%→18.57% and 17.7%→27.14% on CritPt, a frontier-physics benchmark, in two specific paired setups, on a report the authors themselves label a beta technical report. The paper’s claim is that genes can support iterative improvement — not that double-digit lifts arrive everywhere. The controlled probes, not the CritPt endpoints, are the general evidence.
RecapThe four findings, in the paper’s own order
- Documentation-oriented skills are misaligned with test-time control: the useful signal is sparse, concentrated in a narrow procedural slice, and expanding a compact object into fuller documentation can degrade the average.
- Strategy genes are control-oriented objects, not shorter prompts: the benefit emerges with the strategy layer, survives structural perturbation, beats budget-matched fragments, and is diluted by reattached documentation.
- Gene is the better substrate for accumulation: attached failure history works best in a structured, editable carrier, and even then only selectively.
- Failure belongs in compact warnings, and evolution can bank increasingly useful strategy objects. Distilled warnings beat mixed bundles; two CritPt runs show gene-evolved systems pulling far ahead of their paired base models.
Every table above is quoted from the paper. To see how the two representations differ field by field, continue to Skill vs Gene; for the protocol that makes genes evolvable objects, see the GEP page.