From VLM/VLA's to Embodied Agents - Armen Aghajanyan, Perceptron AI
Perceptron argues that VLMs, VLAs, and world models should be unified into a single 'embodied foundation model' that jointly performs perception, reasoning, ...
By Sean WeldonFrom VLM/VLA's to Embodied Agents: Unifying Perception, Reasoning, and Control
Abstract
This synthesis examines the argument, advanced by Perceptron AI, that vision-language models (VLMs), vision-language-action models (VLAs), and world models should be consolidated into a single embodied foundation model performing perception, embodied reasoning, and control jointly rather than as separate paradigms. The analysis draws on Perceptron's Mark1 model and an associated Data Sparse Mixture of Experts (MoE) architecture to identify two structural obstacles to scaling embodied intelligence: signal sparsity in video modeling, where only approximately 2% of the roughly one million visual tokens per hour of video carry usable supervision, and context bloat arising from continuous sensory streams. Evidence indicates a novel data-substitution scaling law - 10x additional video pre-training data offsetting a 10x reduction in teleoperation data - alongside measurable robustness gains against lighting and background perturbation. Implications for robotics data strategy, model architecture selection, and annotation economics are discussed.
1. Introduction
Physical AI systems intended to operate in the real world must simultaneously perceive their surroundings, reason about spatial and causal structure, and produce actions grounded in that reasoning. The prevailing research landscape, however, treats these as separable problems, giving rise to distinct model families: VLMs that consume images, video, and text to produce text; VLAs that extend VLMs with action outputs; and world models that predict future observations. This fragmentation is convenient for benchmarking and modular deployment but is argued here to be an artifact of research history rather than a principled decomposition of the underlying task.
This paper examines the case for an embodied foundation model - a single model that unifies perception, embodied reasoning, and control within one training objective. Key terminology includes Embodied Reasoning (ER) models, VLMs augmented with grounding and spatial reasoning; VLAs, which add action prediction; world models, which output predicted video; and semantic world models, which learn useful representations of the future without explicit rendering. The central thesis is that unifying these capabilities is not merely an engineering convenience but unlocks scaling behaviors and robustness properties unavailable to siloed architectures.
The analysis covers the intellectual lineage of this approach (§2), two principal technical challenges - signal sparsity and context bloat - along with the Mark1 system built to address them (§3), a set of extractable implementation insights (§4), and a discussion of architectural trade-offs and open problems in embodied reasoning (§5).
2. Background and Related Work
The technical precondition for this work is early fusion multimodal modeling, in which heterogeneous modalities are tokenized and interleaved at the input layer rather than combined via late-stage adapters. This line of research, pursued at Meta's Fundamental AI Research (FAIR) lab, established that a single sequence model can be jointly optimized across modalities, providing the foundation for extending the same treatment to action trajectories and control signals.
A parallel strand of robotics research relies on hardcoded percept supervision, exemplified by AI2's Molmo and related VLAs that supervise on explicitly specified quantities such as gripper tip position. Such approaches improve learning efficiency through dense, well-chosen intermediate targets but require manual specification per task and do not scale automatically. Separately, a framing attributed to Google DeepMind (GDM) distinguishes orchestrator models responsible for embodied reasoning and task decomposition from tactile control policies responsible for low-level actuation - a dichotomy that frames the design space between unified and modular architectures examined in this analysis.
3. Core Analysis
3.1 Signal Sparsity in Video Modeling
Video is training-data-rich but supervision-poor. One hour of video yields approximately one million visual tokens, yet only about 2% carry usable ground-truth signal in naive training setups. Common remedies - transcript prediction, synthetic labeling, and dense pixel prediction - each inject a distorted training signal: too sparse, too synthetic, or too indiscriminate, respectively. Hardcoded percepts, such as gripper tip position tracked in Molmo-style VLAs, mitigate this but require per-task manual design and thus fail to generalize automatically across tasks or embodiments. Perceptron reports having developed a proprietary automatic method allowing models to learn which future percepts matter without hardcoded specification, though implementation details remain undisclosed.
3.2 Context Bloat and Data Sparse Mixture of Experts
Always-on cameras and continuous robotic operation require reasoning over token sequences substantially longer and sparser than text. Naive compression approaches, such as patch-wise averaging, achieve up to 10x reduction but are characterized as architecturally unprincipled "hacks" lacking generalizable structure. Perceptron's response is a Data Sparse Mixture of Experts (MoE) architecture, in which a router predicts, at every layer, which tokens merit processing versus which can be skipped. Visualizations of this mechanism show the model naturally allocating compute to high-density, task-relevant regions - zooming into a graph when relevant, or concentrating on tokens labeled "fruit" when instructed to segment fruit - without hardcoded attention priors. This indicates that token relevance can be learned end-to-end rather than engineered.
3.3 Mark1: An Embodied Foundation Model in Practice
Mark1, described as the first embodied foundation model released by Perceptron, was trained on approximately one petabyte of data spanning text, images, video, and trajectories, including desktop use, video games, and robotics. The model is reported to outperform Gemini 3.1 Pro on embodied reasoning benchmarks at roughly 15x lower cost. A notable emergent property is the reframing of classical computer vision tasks, such as object detection, as agentic tasks: the model writes code, zooms into image regions, and adjusts contrast to locate hard-to-detect objects. This reframing enables Mark1 to serve as a cheap, fast robotic data annotator with self-verification, at a fraction of the cost of using larger proprietary models such as Gemini for the same purpose.
3.4 Scaling Laws and Data Substitution
A central empirical finding is a data-substitution scaling law: 10x more general video pre-training data can substitute for 10x less teleoperation data, which costs approximately $100 per hour to collect. While pure VLA/policy training also benefits from additional teleop data, the gains are markedly larger under unified embodied foundation model training. This is demonstrated on multi-step tasks combining perception (reading book titles) and control (sorting books into bins), where pure VLA architectures struggle but unified models succeed by leveraging shared representations across perception and action.
4. Technical Insights
Several implementation-relevant findings emerge from this work. First, the 2% usable-signal ratio in video tokens motivates architectural rather than heuristic solutions to sparse supervision; ad hoc fixes such as synthetic labeling risk introducing systematic bias into training. Second, the Data Sparse MoE router operates across all layers rather than only at input, allowing free-form compute allocation that adapts to task context rather than fixed spatial priors. Third, the reported 10x video-for-teleop substitution ratio holds at current compute scale and should be treated as an empirical observation rather than an asymptotic law; its stability under further scaling remains unverified. Fourth, robustness gains attributed to joint perceptive-control modeling are supported by online training augmentations, including simulated camera occlusion and variable simulated sunlight direction, suggesting that robustness is partly a data-engineering outcome rather than purely an architectural one. Finally, cardinality and spatial relations (left/right, above/below) are noted as underrepresented in internet-scale data, requiring deliberate data distribution design during pretraining - a limitation for any embodied model relying primarily on web-scale corpora.
5. Discussion
The convergence of perception, reasoning, and control into a single model has implications beyond raw performance metrics. The observed data-substitution scaling law suggests that the marginal value of expensive teleoperation data can be partially offset by cheaper, more abundant video pre-training when models are trained jointly - a proposition with direct economic consequences for robotics data collection strategy. This reframes teleop data scarcity as a solvable allocation problem rather than a hard bottleneck, provided the underlying model architecture supports joint learning.
At the same time, embodied reasoning is explicitly acknowledged as an unsolved research problem despite Mark1's reported frontier performance, and the spectrum between end-to-end VLAs and orchestrator-plus-policy systems remains an open architectural question rather than a resolved one. The GDM orchestrator/tactile-policy framing indicates that even proponents of unification recognize scenarios where modular decomposition may retain practical value, particularly for low-latency control loops where a single monolithic model may be inefficient.
Knowledge gaps include the mechanism underlying Perceptron's proprietary automatic percept-learning method, the generalizability of the Data Sparse MoE approach beyond the demonstrated visualizations, and whether the observed scaling law persists under substantially larger compute budgets or across different robot embodiments. These gaps limit the extent to which the reported results can be independently verified or extended by external researchers.
6. Conclusion
This synthesis has examined the case for embodied foundation models as a unification of VLMs, VLAs, and world models, supported by evidence from Perceptron's Mark1 system and its Data Sparse MoE architecture. The two principal technical contributions - an automatic approach to sparse video signal allocation and a learned mechanism for context compression - address structural bottlenecks that have historically constrained video-based training at scale. The reported data-substitution scaling law and robustness improvements suggest practical value for robotics practitioners facing expensive teleoperation costs.
Practical takeaways include treating video pre-training and teleoperation data as substitutable resources within a joint training objective, prioritizing architectural solutions over heuristic compression for long-context video, and designing training data distributions deliberately to cover underrepresented spatial relations. Future work should focus on independent verification of the reported scaling behavior and clarification of when modular orchestrator-policy architectures remain preferable to fully unified models.
Sources
- From VLM/VLA's to Embodied Agents - Armen Aghajanyan, Perceptron AI - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.