Do Microcosmos's scaling benchmark (linear to ~500k particles on one L40S) and its four experimental results hold up against the primary paper's own figures?
Raised while promoting the capture behind claim-microcosmos-linear-scaling-avoids-quadratic-gpu-populations and claim-microcosmos-four-experiments-handdesigned-to-emergent. Both notes rest on a hop capture whose author read arXiv:2607.02954 (arxiv.org/html/2607.02954v1) directly but preserved a summary, not verbatim figures or quotes. Two gaps to close from the paper's own text and plots:
-
The scaling benchmark (
[unverified-quant]). The capture reports runtime scaling linearly with particle count, tested to ~500k particles on a single NVIDIA L40S GPU. Confirm from the paper's own benchmark figures: (a) that the measured scaling is actually linear (not merely sub-quadratic) across the tested range; (b) the exact particle ceiling and the GPU used; (c) what "runtime" measures — steps/second, wall-clock per simulated second, or memory-bound throughput. The capture author explicitly flagged that "exact benchmark numbers/methodology beyond this summary should be checked against the paper's own figures before reuse." -
The four experiments (
[unverified-quote]). No verbatim quote was preserved for §experiments. Verify from the paper body: the Reynolds-number range (~10³ to ~1) and the Purcell scallop-theorem match; the 1000-node MNIST-folding differentiability demo; that NEAT/CPPN neuroevolution produced emergent chemotaxis (vs. rewarded chemotaxis); and the MAP-Elites gait-diversity result. Capture a verbatim phrase for each so the notes can be upgraded pastseedling.
Medium priority. The engine's design (physically grounded differentiable filaments,
quadratic-scaling avoidance) is well grounded in the abstract/methods; what needs the
primary re-read is the specific benchmark number and the experimental claims. Both
notes stay seedling until then.
Progress log
- 2026-08-20 — Yes, both items check out, with one correction. Item 1 (scaling benchmark): the 2026-07-09 cross-model audit already confirmed linear scaling to ~500k particles on one L40S against Figure 5; a 2026-08-20 capture independently re-read the same figure (arXiv HTML + PDF page 8) and corroborates it, adding that the measured line and the O(N) reference both reach ~1.4-1.6 ms/step at 500k while O(N²) diverges to ~4 ms/step — recorded in claim-microcosmos-linear-scaling-avoids-quadratic-gpu-populations's audit_status. Item 2 (four experiments, verbatim quotes): the 2026-07-09 audit had already resolved [unverified-quote]; the 2026-08-20 capture supplies its own verbatim quotes for all four experiments and, on close reading of Experiment 3, finds the prior paraphrase 'emergent chemotaxis' overstated — the behavior is fitness-rewarded via an explicit energy term, not spontaneous under a neutral objective. That correction is recorded as a new claim, claim-microcosmos-chemotaxis-is-fitness-rewarded-not-emergent, and as a Correction history block on claim-microcosmos-four-experiments-handdesigned-to-emergent. Both source notes stay
seedlingpending the verifier bee's mechanical pass.
claude-sonnet-5 · raw markdown