'Skill issue: stop deploying vision language models, use them with Skills - Merve Noyan, Hugging Face'

Developers should stop using vision language models (VLMs) directly for production tasks and instead use them as automated labelers/judges to train small, fa...

By Sean Weldon

Skill Issue: Repositioning Vision Language Models from Inference Engines to Training Infrastructure

Abstract

This synthesis examines a proposed paradigm shift in applied computer vision: rather than deploying vision language models (VLMs) directly for production inference, developers should employ them as automated labelers and judges within a distillation pipeline that trains small, fast, permissively licensed detectors. The methodology - termed "vibe training" - uses Qwen 3.5-9B for pseudo-labeling, an ensemble of two lightweight VLM judges (Gemma 4 E4B and LFM 2.5VL) for quality filtering, and RF-DETR as the student architecture, all orchestrated through a toolkit exposing curated CV models to coding agents. Experiments on road sign detection and document parsing demonstrate competitive mean average precision and evidence of generalization beyond the teacher model's own annotations, at a total pipeline cost of approximately $3-4. Licensing awareness, judge-merging strategy, and the limits of coding-agent common sense emerge as decisive practical concerns for implementers.

1. Introduction

The proliferation of general-purpose VLMs has lowered the barrier to entry for computer vision development, encouraging practitioners to prompt large multimodal models directly for detection, segmentation, and document understanding tasks. This approach fails along three axes relevant to production systems: latency, robustness, and licensing compliance. VLMs do not reliably reach the 30-40 frames per second (fps) required for real-time edge deployment, and specialized architectures such as RF-DETR consistently outperform general-purpose VLMs on narrow detection benchmarks.

A less frequently examined failure mode concerns licensing. The widely deployed YOLO family carries an AGPL 3.0 license, obliging commercial users to either disclose derivative source code or purchase a commercial license. A substantial number of production deployments appear to violate this obligation unknowingly, reflecting a broader pattern in which model selection is driven by accuracy or familiarity rather than legal exposure.

The central thesis advanced here is that VLMs are better positioned upstream in the development lifecycle - as labelers and judges rather than inference engines. This document examines the architecture of such a pipeline (termed "vibe training"), an accompanying toolkit for exposing CV models as agent-callable tools, empirical results from two case studies, and the practical pitfalls encountered when coding agents are given responsibility for dataset construction.

2. Background and Related Work

The work is motivated by an emerging pattern in which large perception models are exposed as tools to a smaller reasoning model, exemplified by configurations pairing SAM 3.1 with Gemma 4. This inverts the conventional hierarchy in which a language model is expected to perform perception directly; instead, the language model orchestrates calls to specialized vision tools.

The proposed WebVision toolkit generalizes this pattern by exposing a curated set of CV models - spanning detection, segmentation, pose estimation, depth estimation, and OCR - as callable tools for coding agents. Model selection follows an explicit, ordered rubric: license (Apache 2.0 or MIT preferred, non-commercial licenses flagged) takes priority over benchmark performance (derived from Hugging Face leaderboards and recent CV conference literature), which in turn precedes informal qualitative assessment ("vibes"). This ordering is a deliberate design stance reflecting the observation that licensing exposure is a more common and consequential production failure than a marginal accuracy deficit.

3. Core Analysis

3.1 Pipeline Architecture and Judge Design

The "vibe training" pipeline comprises three stages: VLM labeling → VLM-as-judge filtering → student model training. Qwen 3.5-9B generates candidate bounding-box annotations over an unlabeled image dataset. Rather than passing raw token-based coordinate outputs to the judging stage, bounding boxes are overlaid visually on the images themselves before being sent to the judge models - a design choice intended to make judgment more analogous to human visual review than to numerical verification.

Two judges evaluate each labeled example: Gemma 4 E4B (8B parameters) and LFM 2.5VL (approximately 2B parameters). Notably, this ensemble of two smaller judges is reported to outperform reliance on a single larger judge. Judgment merging uses a minimum agreement (logical OR) strategy - an example is retained if either judge approves it - rather than a consensus (logical AND) requirement. This choice addresses observed judge imbalance: LFM tends to reject examples more frequently than Gemma, depending on the task, and a consensus requirement would discard an excessive share of usable data. For larger datasets where recall is less critical and dataset volume permits aggressive filtering, a consensus approach is instead recommended.

The filtered pseudo-labeled dataset is then used to train RF-DETR (medium or large variant) as the deployable student model. Because RF-DETR is comparatively small, training is feasible on a single L4 GPU. Infrastructure for the pipeline draws on Hugging Face Jobs for batch labeling and training, Inference Providers for serverless routing of labeling requests, and the Hub for dataset, model, and bucket storage. The reported end-to-end cost of the pipeline is approximately $3-4.

3.2 Empirical Results

The pipeline was evaluated on two tasks. The first, road sign detection, permitted comparison against ground-truth annotations and achieved a favorable mean average precision at IoU 0.50 (mAP50), with an expected performance gap relative to ground-truth-trained models attributable to training on Qwen-generated pseudo-labels rather than human annotations. The second, document parsing, represented a novel task without available ground truth. In this case, the trained RF-DETR model generalized beyond its training distribution, detecting a signature in a test document that Qwen's own original annotation had missed. This result is attributed to the strength of RF-DETR as a backbone architecture rather than to any property of the labeling pipeline itself, suggesting that student model capacity can partially compensate for imperfect teacher supervision.

3.3 Coding Agent Limitations and Human Oversight

Despite automation across labeling and judging, several stages retain a human-in-the-loop requirement. Label descriptions and prompts are auto-generated by coding agents but require human approval before use, since agent-generated prompts are not treated as reliable without review.

A more consequential limitation concerns data augmentation decisions. Coding agents, including capable models such as Opus 4.8, were found to lack domain-specific common sense for computer vision tasks: agents applied horizontal flip augmentation to traffic sign images (inverting text and directional semantics) and color jitter to traffic light images (corrupting the semantic meaning of color in that domain). This necessitated a patch allowing users to selectively disable or enable specific augmentations, since agents could not be trusted to infer which transformations preserve task-relevant semantics.

4. Technical Insights

Several implementation-relevant findings emerge from the pipeline design:

Trade-offs include the continued necessity of human review for prompts and augmentation configuration, and the acknowledged mAP50 gap versus ground-truth training, which represents a quantifiable cost of relying on pseudo-labels rather than human annotation.

5. Discussion

The pipeline described here reframes the role of VLMs in applied computer vision: rather than competing with specialized detectors at inference time, VLMs contribute where their strengths - broad semantic understanding and zero-shot labeling capability - are most applicable, while a compact architecture such as RF-DETR handles the latency- and robustness-sensitive deployment target. This division of labor addresses the three failure modes identified in Section 1 simultaneously: latency is resolved by deploying a small model, robustness is improved by dedicated training rather than zero-shot prompting, and licensing is addressed by prioritizing permissive models both in the teacher-judge stack and in the deployable student.

A notable gap concerns the reliability of coding agents as autonomous pipeline operators. The augmentation failures observed (inappropriate flipping and jittering) indicate that agentic automation of the ML pipeline is not yet trustworthy without domain-specific guardrails, and that human oversight remains necessary at design decision points even when data generation and judging are automated.

The stated future direction of exploring self-improvement loops - training VLMs on their own pipeline outputs - while currently deprioritized in favor of task-specific model deployment, suggests an anticipated trajectory toward tighter integration between the labeling and student model stages, potentially reducing the current dependency on external teacher models altogether.

6. Conclusion

This synthesis presents evidence for repositioning VLMs as labeling and judging infrastructure rather than production inference engines. The "vibe training" pipeline, combining Qwen 3.5-9B labeling, dual small-model judging, and RF-DETR training, achieved competitive detection performance and generalization at minimal cost, while surfacing practical concerns around judge imbalance, licensing, and the limited common sense of coding agents in vision-specific augmentation decisions. For practitioners, the primary takeaways are to treat license terms as a first-order model selection criterion, to prefer judge ensembles with permissive merge strategies over single large judges, and to retain human oversight over prompt and augmentation configuration even within otherwise automated pipelines.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub