Economic outcomes and research behavior
GPT-5.5’s no-notes transfer episode tested four candidates, scoring $3,788, $990, $3,195 and $2,232 in development net active P&L. It retained the first: a quiet-asset reversal without the seed’s stress filter. Softening the low-activity preference, rewarding volatility, and then strengthening the low-activity preference all scored worse. The first change also altered the deterministic tie-breaker, so its gain is not a clean filter ablation. The selected candidate still failed the paired lower-bound and quarter-concentration checks.
GPT-5.5’s source episode ended with four calibration failures and no source notes. It contributes to the completion count, but supplies neither an economic source result nor evidence of knowledge transfer. The frozen transfer payload is empty; failed evaluations are not converted into zero returns.
Sol’s unrelated-note condition retained the original seed at +32.13 points after four candidates scored −34.47, +24.04, +19.70 and +17.67 points. Its difference from the no-notes control is −3.33 points. Both note conditions received empty files: together, these outcomes illustrate run-to-run variation, not the value of useful versus irrelevant knowledge.
Sol’s first source-note condition also selected the stateless reversal formula: +35.69 points, a +0.23-point difference from its no-notes control. The source run wrote no notes, and the selected code differs in its tie-breaker rather than its economic formula. Its four attempts scored +24.39, +35.69, +30.95 and +17.84 points. This is one observed pair, with no estimable research-run confidence interval and no demonstrated knowledge-transfer benefit.
In Sol’s first no-notes transfer run, a liquidity-dependent momentum/reversal hypothesis scored −13.69 percentage points relative to the benchmark. Removing the seed’s lagged stress filter produced the selected candidate at +35.46 points, versus +32.13 for the seed. A half-strength filter scored +30.97; emphasizing extreme losers scored +29.63 and passed the quarter-concentration check, but remained below the selected candidate. All four are adaptive development results. The selected formula also appeared in Sol’s separate refinement run with a different deterministic tie-breaker and a +33.43-point score. This limits attribution of the gain to a new economic insight; neither candidate establishes held-out alpha.
Scores measure the value of executable strategies within the task’s economic assumptions. Research traces show how an agent changes a strategy and responds to another researcher’s evidence.
Observed development case: Astra, repetition 1. Four CORAL islands each received two evaluations. With sharing enabled, two researchers explicitly built on a completed peer strategy. One adapted another island’s local-stress filter; another combined a shared market-stress measure with its own recovery signal. Both follow-ups underperformed their peer reference, by 2.20 and 3.74 percentage points of development excess return.
The best strategy replaced an on/off market-stress switch with the fraction of assets that had fallen at the previous decision. The isolated team independently reached the same best development result. Sharing therefore changed the research path without improving the winning result in this pair. Both selected strategies failed the existing statistical lower-bound and quarterly-concentration checks. This is evidence of knowledge exchange, not evidence of reliable alpha or a held-out collaboration advantage. Inspect the strategy changes, lineage and evaluator references.
Observed development case: Luna, repetition 1. In isolation, all four researchers improved on their own first submission, but only one beat the supplied seed. The best result used a smooth volatility penalty instead of a sharper filter. Its 35.25 pp development result narrowly exceeded the sharing team’s 34.92 pp. Both winners still failed statistical lower-bound and quarterly-concentration checks. Inspect the revisions.
Observed development case: Sonnet, repetition 1. Three of four isolated researchers improved their first submission. The winner changed a volatility penalty into a reduction of the whole signal after severe prior market shocks, raising its development score from 32.74 to 36.09 pp. One island’s flow-filter reversal made its result worse. The winner still failed statistical lower-bound and quarterly-concentration checks. One charged submission used tune mode; an operator invoked native stop after the eight-call budget was exhausted, and two scoreless rejected attempts remain in the audit. Inspect the revisions and closure evidence.
Observed Sonnet sharing revision. In the sharing run, one researcher noticed that replacing a neutral multiplier with a volatility rank also reduced average signal strength. It doubled that rank multiplier to address the mismatch. Its next result rose from 21.84 to 24.84 pp, still below the unchanged seed at 32.13 pp. The code also changed its tie-break salt, so the gain cannot be attributed solely to the multiplier. This is a documented revision of the researcher’s own strategy, not evidence of peer benefit. Inspect the code change and evaluator evidence.
The sharing team selected a different idea: scale the volatility penalty by the fraction of assets that had fallen, instead of an on/off market switch. Its 34.05 pp result trailed isolation’s 36.09 pp by 2.04 pp. Both winners failed statistical lower-bound and quarterly-concentration checks. Sharing used eight charged attempts but only seven received scores: a third submission by one researcher failed before native admission, leaving a 3/2/2/1 allocation. No replacement slot was supplied. The comparison includes this allocation failure and the isolated run’s tune/stop caveat; it does not isolate the effect of sharing at a balanced allocation. Inspect both selections and the allocation audit.
Observed development case: Fable 5.1, refinement repetition 1. Its first change scored 3.27 pp versus the seed’s 32.13 pp. The researcher then revised its public-data fitting method to account for portfolio selection and turnover. Adding a preference for less-traded names scored 21.88 pp; reversing that preference scored −52.58 pp. A final change scored 21.54 pp. All four revisions stayed below the seed, which was retained at every checkpoint: refinement gain 0.00 pp. The trace shows adaptation to feedback, but no improvement in the selected strategy in this repetition. Inspect all submissions, frozen strategies and checkpoints.
Observed development case: Sol, refinement repetition 1. Changing the volatility preference scored 30.18 pp, below the seed’s 32.13 pp. Removing the market-stress volatility filter entirely then scored 33.43 pp, a gain of 1.30 pp. Stronger concentration in less-traded names scored 26.98 pp; combining reversal with momentum scored −13.71 pp. Sol retained the simpler second submission at the final checkpoint. That selected strategy still failed statistical lower-bound and quarterly-concentration checks. Inspect the four submissions and frozen selections.
Observed development case: GPT-5.5, refinement repetition 1. Replacing the on/off market-stress rule with a gradual penalty scored 35.25 pp, up 3.12 pp from the seed. Favoring less-traded names more strongly reduced costs from $1,744 to $1,327 and passed the quarterly-concentration check, but lowered the score to 30.09 pp. Delaying the stress penalty scored 33.33 pp; emphasizing larger prior price falls scored 32.18 pp. The first revision remained selected at every checkpoint. It still failed statistical lower-bound and quarterly-concentration checks. Inspect the four revisions, risk measures and frozen checkpoints.
Observed development case: Opus 5, refinement repetition 1. Favoring fresh price drops scored 19.50 pp; reversing that preference to favor continuing declines scored 20.87 pp. Both stayed below the seed’s 32.13 pp. Opus then abandoned that filter. Avoiding the least-traded names scored 15.01 pp; increasing the preference for less-traded names scored 26.94 pp. The researcher changed its hypothesis in response to feedback, but none of its four revisions improved the selected strategy. The unchanged seed was retained at every checkpoint: refinement gain 0.00 pp. Inspect the revisions, evaluator results and frozen checkpoints.
Observed development case: Haiku 4.5, refinement repetition 1. Replacing lagged market stress with a continuous volatility adjustment scored 20.46 pp, below the seed’s 32.13 pp. A purported threshold change produced the same score: the formula rescales the volatility multiplier instead of changing an activation boundary. The third submission restored the seed’s scoring logic while describing it as a new hybrid, returning to 32.13 pp. Halving the volatility penalty then scored 29.29 pp. The seed remained selected at every checkpoint, giving 0.00 pp refinement gain. These traces distinguish a change in the research narrative from a change in executable behavior. Inspect the threshold formula, the restored seed logic and all frozen results.
Observed source research: Sol, repetition 1. The first submission failed calibration. A revised implementation received a score of −$961.80 in net active P&L on the public July–December 2023 source window; two later ideas scored −$1,982.48 and −$2,138.99. These source scores are separate from the 2024 capability table. No transfer-notes file was produced, so the frozen source-note and unrelated-note conditions both contain empty files. Their later outcomes cannot demonstrate transfer of written knowledge. Inspect the closed source run.
Historical exposure
Withholding test data prevents feedback from that evaluator. It does not exclude market knowledge acquired during pretraining. Dated data audits, controlled tool access and prospective evaluations are needed for stronger temporal claims. Masking names and dates is a partial mitigation, not proof of an uncontaminated task.
Task coverage and uncertainty
Repeated runs measure variability in the research system. Different market periods measure a different source of uncertainty. Many assets or research briefs from the same period cannot substitute for independent market conditions. The reference task count must be revisited after an outcome-free data audit; narrow confidence intervals from only a few periods would be misleading.
Risk and deployment
Net excess return is reported with portfolio return, benchmark return, maximum drawdown, turnover, exposure and concentration. These risk results are not fabricated in this preview and would be required in an empirical release. Historical paper performance does not establish that a strategy can be traded at the assumed costs or capacity.
Scope of the comparison
Native harnesses are part of the evaluated systems. A difference between Sonnet in Claude Code and Astra in Codex cannot be attributed solely to model weights. Collaboration results depend on the sharing policy; transfer results depend on the source tasks and memory construction. These are conditional capability measurements, not universal rankings.