What happened
Hydra has changed its benchmark-diff setup so a pull request and its merge base are measured back-to-back on the same runner. The supplied record says this is intended to cancel cross-machine CPU spread that can make CI benchmark comparisons unstable.
The work also adds hydra_rts_* gauges to hydra-node for allocation, CPU, and garbage-collection activity. Together, these measures are intended to make throughput and cost differences more comparable across CI runs.
Why it matters
Performance claims are more useful when observers can distinguish a code change from variation in the machine running the test. Same-runner pairing and runtime gauges give the Hydra team and outside observers a clearer basis for identifying potential regressions as scaling work continues.
The practical receipt to watch next is benchmark output using the new paired harness alongside the new runtime gauges. Those results can help test later throughput or cost claims rather than treating headline figures as sufficient on their own.
Watch for paired pull-request-versus-merge-base benchmark results and hydra_rts_* allocation, CPU, and garbage-collection gauges in future CI evidence.
Upstream references and independent checks
Digest dated 2026-08-25 · upstream model claude-sonnet-4-6. Direct links are matched to all 2 upstream source IDs.
- 1Bench variance pairing 1 (#2827)Direct upstream source ·
6467c8a747c1ac163141b003369e7337eabbfcf5 - 2Bench variance pairing 2 (#2828)Direct upstream source ·
4eaed8a00c3c27461363af10482ce51ef417c403
This Research brief was generated by Terra from a dated upstream research digest. It has not received the source-by-source human review required for Reviewed analysis. Material limit: The supplied evidence covers CI and measurement tooling only; it provides no new performance figures and does not establish a throughput improvement or protocol change.
