The Dark Arts of Skill Engineering - Paul Bakaus, Renaissance Geek (Impeccable)

Building durable, production-grade AI skills requires moving beyond prompt engineering into 'harness engineering' - extending the coding agent's capabilities w...

By Sean Weldon

The Dark Arts of Skill Engineering: A Synthesis on Harness Engineering for Production AI Skills

Abstract

This synthesis examines the thesis that durable, production-grade AI capabilities require harness engineering - the extension of a coding agent's runtime through scripts, hooks, sub-agent orchestration, persistent memory, and model-specific tuning - rather than prompt engineering alone. Drawing on the development of Impeccable, an open-source design harness spanning Claude Code, Codex, and Gemini, the analysis documents nine engineering techniques for overriding a model's statistical gravity toward generic output. Evidence derives from a custom evaluation harness employing mixture-of-judges design assessment and per-rule ablation testing enabled by unique XML tag identifiers. Findings indicate that prose-based instruction degrades predictably as model capability decreases, that banning outputs relocates rather than eliminates homogeneity, and that deterministic enforcement mechanisms outperform declarative instruction. Practical implications concern skill portability, distribution standards, and the unresolved problem of automated taste evaluation.

1. Introduction

The rapid adoption of agentic coding tools has produced sustained interest in skills: packaged instruction sets that extend an agent's behavior within a specific domain. The dominant mental model treats skills as extended prompts - carefully composed prose intended to steer a model toward desirable output. This synthesis argues that this model is insufficient. As articulated in the source material, "prompting is a spell; harnessing the magic," and reliable behavior emerges only when instruction is coupled to enforcement infrastructure supplied by the harness, the runtime environment in which an agent executes tools, spawns sub-processes, and interacts with external systems.

The central problem is what might be termed statistical gravity: the tendency of a generative model to converge toward the median of its training distribution. In visual design, this manifests as recognizable artifacts - italic serif typography, capitalized hero sections, kicker text, and a beige palette colloquially known as "Claude beige." This attractor is not static. Practitioners who once identified AI-generated interfaces by purple gradients - a pattern traceable to Tailwind's default theme, echoing jQuery UI's earlier "web orange" phenomenon - now identify different markers, since "slop is a moving target."

The central question follows: what engineering mechanisms, beyond prose instruction, reliably shift model output away from its statistical center, and how portable are those mechanisms across heterogeneous harnesses and model capabilities? This analysis covers the origins of Impeccable, the documented limits of prompting, nine identified "dark arts" of skill engineering, the evaluation methodology used to validate them, and the unresolved distribution challenges facing the broader skills ecosystem.

2. Background and Related Work

Impeccable originated as a personal skill called normalize, intended to reconcile Claude-generated interfaces with an existing design system, initially built atop Anthropic's published front-end design skill. Following internal traction, the project expanded into a full design harness operating across Claude Code, Codex, and other runtimes before being open-sourced at impeccable.style.

The empirical foundation for the claims below is a purpose-built Evals harness replicating the Claude Code, Codex, and Gemini SDKs together with their tool surfaces, including browser screenshot capture and interactive question-answering. An LLM simulates user responses during onboarding, enabling automated end-to-end trials. Output quality is assessed by a mixture-of-experts design judge across approximately twenty niches and multiple frontier models, including GPT-5.5, Opus, and Sonnet, establishing a reproducible measurement layer that distinguishes this work from purely anecdotal prompt-tuning practices.

3. Core Analysis

3.1 The Insufficiency of Prose Instruction

A foundational finding is quantitative in character: "even 250 lines of artisanally crafted, beautiful skill prose cannot change this… it's nowhere near enough" to counteract a model's median tendencies. Banning specific outputs - fonts, purple gradients - does not induce creativity but merely relocates the model to the next nearest cluster in latent space, since "a ban only moves the model around inside its own cluster." This establishes prose as a necessary but categorically insufficient mechanism for behavioral control.

3.2 Nine Dark Arts of Skill Engineering

The analysis identifies nine complementary techniques for overriding statistical gravity. First, adversarial argumentation pairs a design-director LLM with a deterministic detector, synthesized by a main thread to avoid self-grading bias - "two blind opinions beat one confident guess." Second, forced divergence uses anti-attractor techniques: shaving and discarding top predictions, blind ranking of large idea batches, or script-seeded randomness (e.g., a color.js file containing over 100 hand-picked palettes). Third, mixture-of-experts routing avoids blurring instructions in a single monolithic skill by loading distinct MD files per task (critique versus polish) and register (brand versus product design).

Fourth, skills are given long-term memory via a persistent project folder (.impeccable) storing critiques and preferences across sessions, enabling compound engineering across multi-session refactors. Fifth, because buried rules are skimmed by weaker models, scripts such as context.mjs inject dynamic, structured JSON instructions via stdout - more reliably followed than static prose, at the cost of breaking prompt caching. Sixth, hooks that fight back use pre-tool-use and post-tool-use guardrails to enforce rules automatically, with weaker models requiring pre-tool-use blocking rather than post-hoc correction.

Seventh, skills should exploit harness-unique capabilities, such as Codex's in-app browser, to build direct feedback loops - Impeccable's live mode uses a server-sent-events poller that halts to signal the model via stdout. Eighth, the "works on my machine" problem requires harness-specific and model-specific compiled builds, since sub-agent spawning, ask-user tools, background job handling, and hook syntax differ substantially across Claude Code, Codex, and Gemini. Ninth, skills must be designed for the weakest model, since each model exhibits unique overfitting tails - Gemini over-animates images, Codex favors gates, rounded borders, and poor letter-spacing - requiring model-specific rule injection and unskippable, loggable checklist gates.

3.3 Evaluation Methodology and the Taste Problem

Validation relies on ablation testing: every rule in Impeccable carries a unique XML tag identifier, permitting systematic removal and re-addition across models to measure impact via the deterministic detection engine. This grants a level of empirical rigor uncommon in skill development, where changes are typically assessed impressionistically. Notably, taste evaluation remains unsolved; judge models are often "maximalist," rating cluttered viewports more favorably, sometimes necessitating inverted judge responses to correct for this bias.

4. Technical Insights

Several implementation-level findings merit isolation. Context.mjs dynamically injects product.md/design.md content and structured JSON fallback instructions on every invocation - a practical solution to prose being skimmed, though one that trades away prompt caching benefits and therefore incurs latency and cost overhead. Codex requires explicit user permission to spawn sub-agents, unlike Claude Code's programmatic spawning, meaning a skill must detect this constraint and either request permission or explicitly disclose a "degraded experience." GPT-5 mini was found to fail at reliably loading certain MD files and does not consistently activate live mode, illustrating that harness capability is not uniform even within a single vendor's model family. Pre-tool-use hooks are established as superior to post-tool-use correction specifically for weaker models, since preventing bad output is more reliable than correcting it after generation. These findings collectively suggest that skill engineering is closer to systems engineering than prompt design.

5. Discussion

These findings reframe skill development as an infrastructure problem rather than a linguistic one. The recurrence of "statistical gravity" as an explanatory concept suggests that any single-model, prose-only approach to behavioral steering has a hard ceiling, regardless of prompt quality - a claim with implications beyond design, potentially extending to any domain where models exhibit convergent, homogenized output.

The distribution challenges documented - absence of an industry standard for skill distribution, caching and update issues in native marketplaces, and the failure of symlinking approaches (e.g., claude.md to agents.md) at scale - indicate that the ecosystem's tooling has not caught up with the complexity required by harness engineering. The necessity of a custom CLI installer and an open pull request proposing multi-harness compilation support suggests this is an active, unresolved area rather than a solved engineering problem.

The taste-evaluation gap represents a genuine knowledge frontier: automated judges' maximalist bias indicates that current LLM-as-judge paradigms are themselves subject to the same statistical gravity they are meant to detect, raising questions about the reliability of any evaluation loop that is not grounded in deterministic, rule-based detection.

6. Conclusion

This synthesis demonstrates that reliable, production-grade AI skill behavior depends on harness engineering - scripts, hooks, sub-agent orchestration, persistent memory, and model-specific compilation - rather than prompt refinement alone. The nine documented techniques, validated through a custom ablation-capable evaluation harness, provide a reproducible framework for practitioners building skills intended to survive across multiple models and harnesses. The practical takeaway is that skill engineers should treat prose as a starting point rather than a solution, and should expect to build enforcement infrastructure, model-specific tuning, and portable distribution tooling as first-class components of any durable skill.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub