A mixed-reality prototype where the user speaks or types a prompt, a generative model makes a matching 3D object, and the object then falls, can be grabbed and thrown, and makes a material-appropriate sound on a Meta Quest 3. The material comes from the words in the prompt, not from any vision model.