SciUniverse

Most science and engineering benchmarks start after the experiment, with clean data ready to analyze. That matters, but it leaves out most of the work: choosing materials, operating instruments, running experiments, debugging failures, and adapting to real constraints.

SciUniverse is a continuous benchmark measuring those capabilities in the physical world. Each task is run or grounded in a model-native research facility in San Francisco and graded on verifiable scientific outcomes.

Loading task space…

Loading results and task space…