Modality Misalignment and Originality Attribution in Short-Form Video - Aditya Gautam, Meta

Detecting modality misalignment and unoriginal content in short-form video at scale (100M+ videos) requires a multi-agent system of specialized, distilled sm...

By Sean Weldon

Modality Misalignment and Originality Attribution in Short-Form Video: A Multi-Agent Systems Approach

Abstract

Short-form video platforms operating at a scale of 100 million or more assets confront two integrity problems: intermodality misalignment, where content diverges mid-stream from what was initially presented, and unoriginal content, enabled by generative tooling that trivializes duplication. This synthesis examines an architecture, drawn from production experience at Meta, that addresses both through a multi-agent system comprising a Reviewer orchestrator, a Perceiver vision-language expert, and a Retriever indexing service. The approach combines domain-adapted small-scale Vision-Language Models (VLMs), knowledge distillation, quantization, and holistic node-level evaluation. Findings indicate that specialization and decomposition outperform monolithic frontier-model deployment on both accuracy and unit economics, with spatial-temporal reduction, similarity-based caching, and metadata pruning serving as primary cost levers. The analysis offers practical guidance for engineers building content-integrity systems at comparable scale.

1. Introduction

Content moderation and integrity systems for user-generated video operate under conditions that diverge sharply from academic benchmark settings. Data arrives at massive scale, is adversarially manipulated, spans multiple languages within a single asset, and shifts distributionally on a monthly cadence as new generative AI tools reach consumers. Critically, no clear ground truth exists for the underlying classification problems, since labels are themselves contested and policy-dependent.

Two failure modes are central to this analysis. Intermodality misalignment occurs when a video's content changes mid-stream - for instance, a sports clip transitioning into unrelated advertising or political content that was not visible when the user initially engaged. Detecting this requires frame-level and clip-level video understanding capable of localizing temporal anomalies, rather than assigning a single video-level label. Unoriginal content, the second problem, arises because AI tooling makes duplication and transformation of existing uploads inexpensive. Its harms are systemic: attribution is misallocated to re-uploaders rather than original creators, and users experience fatigue from repetitive feed content.

The central thesis is that these problems are tractable at production scale only through a combination of task decomposition across specialized agents, small-scale VLMs adapted to in-house data distributions, evaluation instrumented across the entire agent graph rather than only the terminal output, and adaptive optimization that allocates compute proportional to problem difficulty. Sections 2 and 3 establish the theoretical and architectural context; Sections 4 and 5 extract implementation guidance and discuss open questions for the field.

2. Background and Related Work

Frontier VLMs are pre-trained predominantly on curated web corpora, which biases them toward clean, well-composed imagery and standard language use. This inductive bias maps poorly onto user-generated short-form video, which is artifact-laden, contains overlaid multilingual text, and is frequently low-bitrate or intentionally manipulated. The design philosophy underlying this work is deliberately narrow in scope: general-purpose capability is treated as irrelevant to the objective function, as reflected in the observation that "I really don't care if the model can solve a coding problem. I only care about my domain specific problem."

A second foundational commitment concerns architectural restraint regarding multi-agent design. Multi-agent systems introduce orchestration overhead, additional latency, and expanded failure surface, and are justified only when the task genuinely requires heterogeneous capabilities: "If you can solve a problem with one agent, one LLM, we don't need to actually do it on multi-agent systems." In this domain, the justification rests on the observation that the problem decomposes cleanly into retrieval, perception, and reasoning - three capabilities with distinct data requirements, latency profiles, and hardware needs. This connects to a broader methodological toolkit including instruction fine-tuning via a projector layer bridging vision and language modalities, Direct Preference Optimization (DPO) for continuous improvement, and knowledge distillation for cost control.

3. Core Analysis

3.1 Multi-Agent Architecture Design

The system is structured around three specialized agents. The Reviewer agent functions as a centralized orchestrator and API gateway, performing signal decomposition and temporal analysis across incoming data streams. The Perceiver agent is a sophisticated VLM expert that fetches raw video, decomposes it into clips using semantic embeddings and temporal change detection, and compresses visually similar frames to reduce redundant processing. Its output is structured JSON containing clip-level embeddings, tags, OCR extractions, natural language descriptions, and timestamps. The Retriever agent maintains adaptive metadata indexes - an inverted index for topics, a vector database for embeddings, and a graph database for entities - and performs offline clustering on distributed compute (e.g., a Ray cluster) to precompute cluster identifiers used during online inference.

This decomposition allows the Reviewer to detect anomalies, such as sports content shifting into political content, by analyzing the Perceiver's temporal output, then query the Retriever to confirm whether similar clips exist elsewhere in the corpus, indicating either misalignment or duplication. Online retrieval further reranks candidates using spam and quality classifier scores, and the Reviewer incorporates real-time behavioral signals - likes, dislikes, comments, and sentiment - into its final determination. This layered structure reflects the guiding principle that "decomposition is key," applied only "as and when necessary."

3.2 Domain-Specialized Model Development

Because frontier VLMs are trained on clean web data, they underperform on messy, user-generated content without adaptation. The described pipeline pre-trains vision transformers on domain-specific image tokens and language when this produces a measurable performance delta over standard fine-tuning alone. Instruction fine-tuning then bridges the vision encoder and language model via a projector layer, trained on domain-specific JSON schema outputs that encode role, policy, tool-use, grounding metadata, and chain-of-thought reasoning traces.

Continuous improvement is achieved through a DPO phase that samples production data daily, evaluates it via LLM-as-judge, and routes uncertain cases to a human review queue. A human-in-the-loop process traces failures across chain-of-thought reasoning, tool calls, and retrieval steps to identify root causes, feeding positive and negative samples back into a continuous retraining loop that responds to production drift.

3.3 Cost Optimization and Scalability Techniques

Given that frontier-size VLMs are not economically viable for narrow, internal-only problems at 100M+ scale, the system employs both off-policy and on-policy knowledge distillation to produce smaller specialized models, alongside quantization experiments across float variants to balance performance retention against compute savings. Comparison tables across distillation and quantization configurations identify optimal cost-performance tradeoffs, typically targeting a performance retention threshold (e.g., 95% of baseline).

Three scalability techniques reduce processing load independent of model size: spatial-temporal reduction, which compresses similar frames rather than sampling at a fixed frame rate; caching, which uses high similarity scores to skip redundant multi-agent processing for viral or duplicate content; and metadata pruning, which filters out videos from high-authenticity creators or already well-performing content before they enter the full pipeline.

4. Technical Insights

Several implementation-level findings merit emphasis for practitioners. First, temporal change detection - rather than fixed-interval frame sampling - is used to compress video into representative frames, reducing both storage and inference cost while preserving anomaly-relevant signal. Second, the three-tier retrieval index (inverted, vector, graph) allows the Retriever to serve heterogeneous query types (topical, semantic, relational) without maintaining separate infrastructure stacks for each. Third, offline clustering on distributed compute decouples expensive batch computation from latency-sensitive online inference, a pattern applicable broadly to retrieval-augmented systems operating at scale.

A key trade-off concerns pre-training versus fine-tuning: pre-training vision transformers from scratch on in-house data is only justified when it yields a measurable delta over fine-tuning, implying that teams should benchmark this decision rather than default to either approach. Similarly, distillation and quantization decisions are made empirically via comparison tables rather than fixed heuristics, suggesting that cost-performance optimization is treated as a continuous experimental process rather than a one-time architectural choice. Finally, the reliance on LLM-as-judge introduces a monitoring requirement: judge performance must itself be tracked for drift against human review queues, with retraining triggered when divergence is detected - an often-overlooked meta-evaluation layer.

5. Discussion

These findings suggest that production-scale content integrity systems increasingly resemble distributed software systems more than monolithic models, with the attendant engineering discipline of instrumentation, caching, and load-shedding applied to AI reasoning pipelines. The emphasis on holistic evaluation - extending precision, recall, and F1 metrics beyond the end task to retrieval quality, latency, chain-of-thought efficiency, and robustness at each node - reflects a broader industry shift toward treating agentic systems as pipelines requiring per-stage observability rather than black boxes evaluated solely on final output.

An open question concerns generalizability: the described architecture is tailored to video-specific problems (temporal misalignment, duplication) and it remains unclear how directly this decomposition transfers to other modalities or integrity problems with different failure geometries. Additionally, the absence of clear ground truth for misalignment and originality labels raises questions about how evaluation metrics themselves are validated, since precision and recall calculations presuppose a trusted reference standard that the source material acknowledges does not fully exist.

6. Conclusion

This analysis has described a multi-agent architecture - Reviewer, Perceiver, and Retriever - designed to address intermodality misalignment and unoriginal content at the scale of 100 million or more videos. Its core contributions lie in demonstrating that task decomposition, domain-specialized small-scale VLMs, and holistic node-level evaluation collectively outperform monolithic frontier-model deployment on both accuracy and cost. Practitioners building comparable systems should prioritize empirical validation of pre-training and distillation decisions, instrument every pipeline node rather than only terminal outputs, and treat evaluation infrastructure itself as a first-class system component requiring continuous drift monitoring.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub