MirroS releases AgentGarten — a framework for training AI agents in simulated code worlds
Read more
Pandaily
pandaily.com

MirroS releases AgentGarten — a framework for training AI agents in simulated code worlds

The research team at MirroS has introduced AgentGarten, a framework that provides AI agents with an environment for interaction, observation, and learning. According to a report from QbitAI on October 9th, all necessary code, the technical report, and the project page are publicly available.

AgentGarten divides the world simulation task into two parts. The code handles the game rules, physics, contacts, and layout, ensuring an explicit state and predictability of every outcome. Subsequently, a neural renderer transforms the schematic geometric drawing exported by the code (such as depth or surface normals data from the agent's camera) into a realistic next frame.

The team asserts that this approach avoids two common compromises: scenes created with game engines, which require expensive artistic refinement to achieve realism; and world video models whose physics are hidden within the neural network and cannot be queried or modified.

The renderer generates video in short streaming segments, allowing the agent to act, see the result, and make the next decision. MirroS trained the system using a method they call Adversarial Forcing. This method involves precise reproduction so that gradients reach the model-generated history, alongside a discriminator trained on real video to prevent drift and grid artifacts during long sequences of actions.

Thanks to the use of custom Triton kernels, CUDA graphs, and a lightweight decoder, the system supports a frame rate exceeding 30 frames per second at 480p resolution on a single GPU.

To test the training process, the team reused the hide-and-seek game from a 2019 study. In that study, reinforcement learning required about 25 million episodes before the hiders built shelters, and about 100 million before the seekers started using ramps. In AgentGarten, agents only saw first-person rendered frames and controlled actions via short Python scripts. After each round, they analyzed their gameplay sessions and recorded lessons in a 'rulebook' passed to the next round. The hiders began building shelters by the fourth round, and the seekers could overcome walls using a ramp by the tenth round.

A similar cycle was launched in four other environments over four rounds each. Performance metrics improved: the companion dog's evaluation increased from 13 to 19, the time taken for two cars to cross a bridge decreased from 71 to 41 seconds, two shepherd dogs managed to gather all four sheep instead of three, and a loader for a quarry completed its task 31 seconds ahead of schedule.

MirroS views this work as a step toward creating agents that continuously improve in physical conditions, emphasizing that it is a research environment, not a ready-to-deploy robot product.

Popular