Things to improve:
Have a flywheel vlm-in-loop eval
Create 5 possibilities and pick the best one
Run a bunch of cases and pick up only the successful cases with RAG
Endgame, could also do a council-based approach
https://tml.stanford.edu/
For sims
https://arxiv.org/pdf/2509.26633
https://jiaxin-lu.github.io/humoto/
Potential idea: use COSMOS (or a more egocentric version) to fetch a video (or get a very specified prompt w/ world model to generate such a thing?), convert to sim for context and then animate with the in-context example (basically sim2real2sim)
e.g. dataset https://signavatars.github.io/