AI Inside VR: Adaptive Training Scenarios and Virtual Instructors
VR made training safe and repeatable. AI makes it personal. Combining language models, computer vision and simulation lets a trainer adapt to each learner, explain mistakes in plain language and generate new scenarios on demand. Here is how we are building it, and what still needs a human.
By Solomiia Kots
CEO at Mirko

For years a VR simulator had one difficulty curve and one script. Every trainee saw the same forklift aisle, the same leak, the same checklist. The simulators we build now watch what the learner does, decide what they need next and explain why, using the same agent harnesses we deploy in business software. The headset became an interface to an AI tutor.
Adaptive scenarios
The scenario graph that drives a simulator is data, which means an agent can traverse it. We track each trainee's error patterns, hesitation and tool choices, and an orchestration layer picks the next scenario to target the weakest skill: more emergency branches for the operator who freezes, more sequencing drills for the one who skips steps. Difficulty adapts within a session, not at the end of a course.
A virtual instructor that can actually explain
The most requested feature from trainers was simple: when a trainee makes a mistake, tell them why in their own words. We connect the simulation state to a language model with a tightly scoped context: the procedure, the current state, the mistake and the safety rationale from the client's documentation. The instructor speaks through text-to-speech, answers follow-up questions and stays inside the procedure, because the harness gives it nothing else to talk about. The same context-engineering rules we use for enterprise copilots apply here.

Context Engineering: Managing AI Attention in Production
Guardrails we put around the virtual instructor
- It explains and asks; it never grades. Scoring stays deterministic from the scenario graph.
- It only cites the client's own procedures and safety documents, with the source shown to trainers.
- Every instructor conversation is logged and sampled for review, like any production agent.
- Trainers can override, correct and add explanations, which feed the evaluation set.
Computer vision for AR guidance
In augmented reality the instructor needs to see what the technician sees. Computer vision recognises the machine state through the headset or phone camera, so AR instructions anchor to the real component and the system knows when a step is actually done. The retail shelf recognition and industrial inspection models we have built transfer directly to this use case.
Generating scenarios
Writing a new scenario used to take a developer a week. With structured scenario definitions, an agent can draft a new branch from an incident report, a trainer reviews it in an authoring tool and it ships the same day. We keep humans in that loop deliberately: generated content is reviewed before any trainee sees it, following the same approval gates we use elsewhere.





