Hi SimplerEnv team,
I maintain OmniSim, an Apache-2.0 simulator built for agent-driven world authoring, execution, inspection, and reproducible evaluation over HTTP/MCP.
SimplerEnv's MMRV and Pearson-correlation framing is exactly the kind of evaluation discipline we want. We would like to implement one small SimplerEnv task in OmniSim and evaluate it through your existing methodology:
- you nominate one Google Robot or WidowX task with a stable observation/action/success contract;
- we handle the robot/task port and publish all scripts and asset-provenance notes;
- we run the same policy set where technically feasible and feed OmniSim's results into the existing MMRV/Pearson tooling;
- we report failures and mismatches, not just successful screenshots.
OmniSim is not photorealistic, so we would avoid presenting this as a visual-matching replacement. The useful question is whether a second simulator can preserve the task contract and policy ranking.
Which single task and policy pair would make the smallest credible pilot?
Best,
Ahmed Fetouh — OmniLink
Disclosure: this note was drafted with the same AI-agent workflow used to build OmniSim and reviewed/sent by me.
Hi SimplerEnv team,
I maintain OmniSim, an Apache-2.0 simulator built for agent-driven world authoring, execution, inspection, and reproducible evaluation over HTTP/MCP.
SimplerEnv's MMRV and Pearson-correlation framing is exactly the kind of evaluation discipline we want. We would like to implement one small SimplerEnv task in OmniSim and evaluate it through your existing methodology:
OmniSim is not photorealistic, so we would avoid presenting this as a visual-matching replacement. The useful question is whether a second simulator can preserve the task contract and policy ranking.
Which single task and policy pair would make the smallest credible pilot?
Best,
Ahmed Fetouh — OmniLink
Disclosure: this note was drafted with the same AI-agent workflow used to build OmniSim and reviewed/sent by me.