SciUniverse
Most science and engineering benchmarks start after the experiment, with clean data ready to analyze. That matters, but it leaves out most of the work: choosing materials, operating instruments, running experiments, debugging failures, and adapting to real constraints.
SciUniverse is a continuous benchmark measuring those capabilities in the physical world. Each task is run or grounded in a model-native research facility in San Francisco and graded on verifiable scientific outcomes.
Loading task space…
SELECTED TASK
Model results
Eval ↓ · Run timeSELECTED RUN
Run milestones
Loading camera…
OPERATOR POVCurrent action
Loading instruments…
INSTRUMENT VIEWPhysical context
Loading results and task space…