Back to AI Research

AI Research

Pixels to Keys: Exploring Spatial and Motion Cues i... | AI Research

Key Takeaways

  • Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics investigates whether an AI model can infer a player’s keyboard inputs from gam...
  • Video games offer scalable environments for studying perception and control in embodied agents.
  • Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training.
  • Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames.
  • Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics.
Paper AbstractExpand

Video games offer scalable environments for studying perception and control in embodied agents. Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as $F_1^{macro}$. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.

Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics investigates whether an AI model can infer a player’s keyboard inputs from gameplay video alone. This is an inverse dynamics problem: instead of predicting what frames an action will produce, the model works backward from frames to estimate the action that occurred. In the paper, the authors study which visual signals, architectures, and training objectives help recover individual key presses when labeled gameplay data is limited.

What the paper is trying to recover

Online gameplay videos are plentiful, but they rarely include the keyboard and mouse inputs used to create them. Manually adding those labels is expensive, motivating inverse dynamics models (IDMs) that can automatically annotate video for later use in behavioral cloning, reinforcement learning, or other embodied-agent systems.
The task is difficult even in a small action space. The same visual outcome can sometimes result from different inputs, while some inputs have weak or delayed visual effects. In a racing game, for example, inertia can continue moving the camera or car after a key is released. Camera motion can also be confused with character motion. Finally, key presses are highly imbalanced: common actions such as moving forward may appear far more often than turning, braking, or pressing a special key.
The study focuses on five navigation keys—W, A, S, D, and Space—and evaluates models on two games. Trackmania provides a third-person driving setting, while Cyberpunk 2077 provides a first-person navigation setting. Both datasets use 20 frames per second and 512×512 images, with separate recordings for training, validation, and testing.

How the models use frames and motion

The baseline input is a five-frame clip, covering 0.25 seconds. The model predicts each key independently, allowing multiple keys to be active at once. The authors also test models that predict the key state for every frame in the clip, rather than only the center frame, but increasing the clip to 16 frames does not improve performance. The final experiments therefore use five-frame inputs and center-frame predictions for the strongest architecture.
The study compares three broad model types: a small convolutional neural network, a video Vision Transformer, and a hybrid Transformer. The hybrid model first uses convolutional layers to extract and downsample local visual features, then applies Transformer blocks and learned action-specific queries to estimate the five key probabilities.
A central experiment adds optical flow to the RGB frames. The authors compute motion fields with RAFT, representing horizontal and vertical displacement between each adjacent pair of frames. This gives the model an explicit motion signal that may help distinguish camera movement, character movement, and other changes in the scene.
The strongest training recipe also uses supervised contrastive pretraining. Two visually perturbed versions of a sequence are encouraged to have similar embeddings when they share the same action label, while sequences with different labels are pushed apart. The model is then trained with a soft-F1 loss, which directly averages performance across keys instead of allowing frequent inactive labels to dominate the objective. This focus on balanced key recovery resembles concerns in action-discriminative world models, although that work studies whether a world model can distinguish the consequences of alternative actions rather than reconstructing keyboard inputs from observed frames.

Results that stand out

On Trackmania, architecture makes a large difference. The CNN reaches a macro-F1 of 0.367, while the video Transformer reaches 0.637 under otherwise similar settings. The hybrid RGB model reaches 0.746 macro-F1, showing that combining convolutional feature extraction with Transformer processing is substantially more effective in this setting.
Adding optical flow produces a more uneven result. The hybrid RGB-flow model has a macro-F1 of 0.791 for W, 0.920 for A, 0.829 for S, 0.655 for D, and 0.822 for Space, compared with 0.746 overall for the RGB-only hybrid model. The most complete Trackmania model adds contrastive pretraining and reaches 0.895 macro-F1, with a micro-F1 of 0.938 and overall accuracy of 0.938.
The per-key results are important because aggregate metrics can hide failures. In the CNN experiment, for example, F1 is 0.902 for A but only 0.075 for D and 0.008 for Space. The paper argues that accuracy is especially misleading when inactive labels and common actions dominate. Macro-F1, which gives each key equal weight, is therefore used as the primary evaluation metric.
Applying the Trackmania architecture and training recipe directly to Cyberpunk 2077 exposes a substantial domain gap. On the five shared keys, the model records 0.715 macro-F1 and 0.617 micro-F1. Its per-key F1 scores range from 0.473 for Space to 0.786 for A. The same approach therefore does not perform uniformly across game mechanics or viewpoints, even when the nominal action set is kept the same.

What the findings leave unresolved

The authors’ failure analysis points to problems that architecture changes alone may not solve. Visual evidence can be ambiguous when camera motion and character motion are entangled. Delayed effects and inertia mean that the frame immediately around an action may not show its consequences clearly. Some actions are also difficult to observe—for instance, turning while a car is pressed against a wall.
These findings suggest several directions for future IDMs: explicitly modeling 3D scene structure, maintaining longer-term state, and using losses designed for imbalanced multilabel actions. The attention analysis further indicates that action queries need access to the parts of the scene that reveal motion, while also being able to separate camera movement from movement caused by the player.
The paper’s broader lesson is that recovering “what happened” from gameplay video is not adequately captured by one overall accuracy number. A model can appear successful while failing on rare but strategically important inputs. More balanced evaluation and per-action analysis are necessary if IDMs are to produce reliable annotations for embodied agents such as the conversational PUBG teammate, whose control system likewise must connect visual perception with discrete game actions.

Comments (0)

No comments yet

Be the first to share your thoughts!