Mini Models Are the Future of Applied AI.
Specialized, low compute systems will turn narrow business workflows into reliable products, and may capture more economic value than the general models beneath them.
Abstract
Frontier models will remain essential for difficult, ambiguous, and multiple step work. But most commercial activity is repetitive, bounded, and specific to a workflow. My thesis is that the largest applied AI profit pool will form around “mini models”: specialized systems that use the least expensive intelligence capable of completing a narrow job, then combine context, tools, rules, verification, and escalation to deliver a finished outcome.
The working thesis
Convenience, not raw intelligence, is what most customers ultimately buy.
People do not want to manage prompts, models, context windows, or tool calls. They want the intake completed, the campaign written, the funnel assembled, and the project system cleaned up. AI creates value when the complexity disappears behind a dependable result.
The winning product may therefore look less like one universal intelligence and more like a set of Lego pieces: narrow capabilities, assembled around a specific job, that simply deliver the magic.
What I mean by a mini model
A mini model is a specialized AI product system, not merely a model with fewer parameters.
It uses the smallest and least expensive model that can reliably complete a bounded task. That model may be small, distilled, fine tuned, locally deployed, or accessed through an inexpensive API. The system around it matters as much as the model itself.
- Narrow objective
- One clearly defined job with an observable finish line.
- Curated context
- Only the business rules, examples, and data needed for that job.
- Tools and actions
- APIs, databases, and software controls that move work forward.
- Verification
- Tests, rules, and human review where errors are costly.
- Escalation
- Uncertain or novel cases route to a frontier model or a person.
Why this becomes possible now
Capability is moving down the cost curve while software orchestration is getting better.
Compact models are becoming capable enough to handle useful language, reasoning, multimodal, and function calling tasks. Knowledge distillation can transfer behavior from larger systems into smaller ones, while model cascades can route only the hardest cases to expensive intelligence.
The result is a new design choice: use a frontier model everywhere, or build a system that spends intelligence only where the task requires it.
Technical referencesGoogle FunctionGemma model card ↗Phi 4 Mini technical report ↗Knowledge distillation ↗FrugalGPT model cascades ↗
The Lego system
The model is one component in a chain that turns intent into an outcome.
A user should not feel this architecture. They should feel speed, consistency, and relief. The best implementation hides the Lego pieces and exposes one simple promise: give the system what it needs, and the work comes back finished.
Business use cases
The strongest early markets are repetitive workflows with abundant context and a measurable output.
Law firm intake
Collect facts, detect missing information, categorize the matter, prepare a summary, and route the prospect to the correct next step, without presenting the system as legal advice.
Email copywriting
Combine brand voice, customer segment, offer, inventory, campaign history, and compliance rules to produce focused drafts and testable variants.
Landing pages and funnels
Turn an offer and audience into page structure, copy blocks, forms, follow up sequences, and experiments connected to conversion data.
Routines and sweeps
Inspect the software stack, reconcile stale work, surface blockers, summarize risk, and propose actions for human approval.
The economics of the edge
Value accrues where AI touches a real workflow, controls an outcome, and runs at scale with low unit cost.
Mini models can lower inference expense, reduce latency, support local or private deployment, and make performance easier to test. But low model cost alone is not enough. Integration, supervision, maintenance, and failure handling remain part of COGS.
The 80% profit thesis depends on application companies owning distribution, workflow context, evaluation data, and the customer relationship. If those advantages remain at the application layer, the specialized system, not the underlying model, can capture the larger share of the economic surplus.
Frontier models still matter
This is a routing thesis, not an argument that large general models disappear.
Frontier models remain the right choice for ambiguous strategy, novel problems, reasoning across multiple domains, complex context workflows, and expert users who can exploit broad capability. They also help design, train, evaluate, and supervise specialized systems.
| Dimension | Mini model system | Frontier model |
|---|---|---|
| Best task | Bounded and repeated | Novel and ambiguous |
| Context | Curated and proprietary | Broad and flexible |
| Economics | Low cost at high volume | Higher cost for hard cases |
| Evaluation | Narrow and measurable | Open ended and harder to score |
| Role | Default execution layer | Escalation and expert layer |
The product and venture playbook
Win a narrow workflow, prove the economics, then add adjacent pieces.
- Choose a painful repetitionFrequent, expensive, and easy to recognize when complete.
- Own the contextBring together the rules, examples, history, and data unique to the customer.
- Define the evaluationMeasure quality, completion, correction rate, latency, and cost per outcome.
- Use the cheapest passing modelSpend capability according to the difficulty and consequence of the task.
- Keep an escalation pathRoute uncertainty to a stronger model or a qualified person.
- Price the resultAlign revenue with time saved, conversion created, or work completed.
- Expand like LegoAdd adjacent workflows only after the first one is dependable.
Risks and falsifiers
A useful thesis must explain what could prove it wrong.
The model may be only a small part of total cost. Integration and maintenance can overwhelm inference savings. Narrow systems can become brittle when the workflow changes. Customers may prefer one general interface, and frontier prices may fall quickly enough to erase the economic advantage of specialization.
The thesis weakens if general models become cheap, fast, private, and reliable enough to handle nearly every bounded workflow without specialized data, evaluation, or orchestration. It also weakens if application companies cannot retain distribution or if model providers absorb the customer relationship.
The response is not blind specialization. It is continuous evaluation: compare the mini system against the frontier baseline, track corrections and failures, and replace the architecture when the economics no longer justify it.
Conclusion