Do Microcosmos's scaling benchmark (linear to ~500k particles on one L40S) and its four experimental results hold up against the primary paper's own figures?
Raised while promoting the capture behind claim-microcosmos-linear-scaling-avoids-quadratic-gpu-populations and claim-microcosmos-four-experiments-handdesigned-to-emergent. Both notes rest on a hop capture whose author read arXiv:2607.02954 (arxiv.org/html/2607.02954v1) directly but preserved a summary, not verbatim figures or quotes. Two gaps to close from the paper's own text and plots:
-
The scaling benchmark (
[unverified-quant]). The capture reports runtime scaling linearly with particle count, tested to ~500k particles on a single NVIDIA L40S GPU. Confirm from the paper's own benchmark figures: (a) that the measured scaling is actually linear (not merely sub-quadratic) across the tested range; (b) the exact particle ceiling and the GPU used; (c) what "runtime" measures — steps/second, wall-clock per simulated second, or memory-bound throughput. The capture author explicitly flagged that "exact benchmark numbers/methodology beyond this summary should be checked against the paper's own figures before reuse." -
The four experiments (
[unverified-quote]). No verbatim quote was preserved for §experiments. Verify from the paper body: the Reynolds-number range (~10³ to ~1) and the Purcell scallop-theorem match; the 1000-node MNIST-folding differentiability demo; that NEAT/CPPN neuroevolution produced emergent chemotaxis (vs. rewarded chemotaxis); and the MAP-Elites gait-diversity result. Capture a verbatim phrase for each so the notes can be upgraded pastseedling.
Medium priority. The engine's design (physically grounded differentiable filaments,
quadratic-scaling avoidance) is well grounded in the abstract/methods; what needs the
primary re-read is the specific benchmark number and the experimental claims. Both
notes stay seedling until then.