An annotated reading guide · compiled 6 October 2026 arXiv:2604.15097 · cs.SE
Skill → Gene

Reading notes on From Procedural Skills to Strategy Genes — how LLM agents should encode experience so it actually controls behavior.

arXiv:2604.15097
Wang · Ren · Zhang
Submitted 16 Apr 2026

Benchmark · the setting

Gene-Bench: where the 4,590 trials happen

Gene-Bench is the name the paper gives its benchmark setting: 45 scientific code-solving scenarios on which every condition in arXiv:2604.15097 is evaluated, for 4,590 retained trials in total. Each scenario asks the model to write a Python program; the program runs in a sandbox; a scenario-specific test script scores it against a predefined set of checkpoints. Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package.

The domain spread is deliberately wide: bioinformatics, neuroscience, chemistry, seismology, climate and atmospheric science, signal processing, network analysis, finance, robotics, and quantum computing. Task shapes vary just as much: structured parsing and transformation, analysis pipelines, simulation and fitting, algorithmic planning, structured output generation. That breadth is what licenses the paper’s representational claim beyond any single niche.

Worked examplesFour scenarios the paper puts on the table

The paper details four scenarios enough to picture the whole bench. They are worth quoting because they show why checkpoint scoring is not a technicality: these tasks have many independent ways to be partially right.

Table 1. Representative scenarios (paper §3.3.2, §B.2).
ScenarioDomainCheckpointsWhat the program must do
S012_uv_spectroscopyChemistry / signal processing—Read UV-Vis spectra from CSVs, detect peaks, compute wavelength, height, FWHM and area, identify dominant peaks, write structured output. The paper’s running example.
S002_spike_behaviorNeuroscience12Read spike times and velocity from a MATLAB .mat file, filter successful trials, bin spikes, align behavior by interpolation, write trial-structured HDF5 with quality-control flags.
S026_earthquake_catalogSeismology14Aftershock identification, completeness magnitude, Gutenberg–Richter b-value via the Aki maximum-likelihood formula, structured CSV/JSON summaries.
S114_obstacle_avoidanceRobotics11Build a valid 2D path via RRT or potential fields, geometric collision checking, path smoothing, and metrics: length, waypoints, clearance, smoothness.

Checkpoint counts as stated in the paper; S012’s count is not given in the sections we annotate, so none is invented here.

Metric & protocolPartial credit, one baseline, fixed decoding

The score is checkpoint-based pass rate. For each trial, the program earns the fraction of the scenario’s checkpoints it passes; a condition’s reported figure is the average over its trials. Complete passes (every checkpoint green) are recorded but treated as a secondary descriptive statistic — the paper’s reasoning is that scenarios differ widely in checkpoint granularity, and binary success would let coarse-grained scenarios dominate the aggregate. Δ values elsewhere on this site are percentage points against the no-guidance baseline within each experiment.

Everything that could confound the comparison is pinned. Two fixed models — Gemini 3.1 Pro Preview and Gemini 3.1 Flash Lite Preview. Low-temperature decoding. A 16,384-token output budget. The task description always enters through the user-facing contents field; the control representation — gene, skill, or nothing — enters separately through systemInstruction. The execution pipeline, timeout policy, and scenario set are shared. The representation is the only moving part.

Table 2. The probe budgets, as reported.
ProbeRetained trialsFocus
Skill Probe1,440Why documentation packages misfire as control
Gene Probe1,890Construction, robustness, dilution, composition
Evolution Probe1,260Carriers, structure, and failure encoding over time
Total4,590—

Counts from §4.1–§4.3 of the paper; Appendix B documents the full protocol.

Two honest limits, both stated in the paper itself: the scenarios are scientific code-solving tasks, so transfer to other agent environments is a hypothesis, not a finding; and the CritPt evolution runs are a separate benchmark with its own pairing. Where those caveats matter to a number, the findings page repeats them in place.

9 / 14 checkpoints pass → trial pass rate = 9/14 ≈ 0.643

Why not score pass/fail? Because these scenarios fail partially and informatively. A program can parse the catalog correctly and compute distances correctly, yet still miss the b-value estimate. Checkpoint scoring keeps that signal: each cell above is one sub-task, and a trial’s score is simply the fraction that lit up.

Illustrative trial, not from the paper’s runs. Scenario size matches S026_earthquake_catalog (14 checkpoints).

FIGURE 1. How checkpoint-based pass rate works. A condition’s reported figure is the average of such trial-level fractions, and a “complete pass” (all cells lit) is recorded only as a secondary statistic — see the findings page for the reported numbers.