[Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation.
-
Updated
Sep 24, 2026 - Python
[Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation.
[Public preview] An advanced AI Benchmark for math, produced by mathmo at St John's College, Cambridge.
[Public preview] The first public quant trading benchmark scored against what professional traders actually made on the same asset
[Public preview] Hard human-written reasoning puzzles: 749 answered items in 15 topics (707 scored core), with protocol and scorer.
To associate your repository with the simreal topic, visit your repo's landing page and select "manage topics."