AI Business Automation Meets the Zero-Token Model
48.57 is the number that matters in Liquid AI’s October 7, 2026 release: that is the reported Decision Index v0.2.1 score for d1-3B, a multimodal decision model built for structured judgments rather than chat. For AI business automation, that matters because the market is starting to split between systems that generate language and systems that simply decide. The latter can be easier to wire into production routing, moderation, inspection, and triage flows where latency and parsing overhead matter more than eloquence. According to a MarkTechPost report on the launch, both checkpoints are already available on Hugging Face and supported in common deployment stacks.
The headline trend: AI business automation is shifting from generation to decision layers
The important shift in this release is not that another open model arrived. It is that Liquid AI packaged a model family that returns structured answers in one forward pass with zero output tokens. Instead of generating a sentence and forcing downstream systems to parse it, the d1 models answer named questions with probabilities, labels, or ordered scores.
That design fits a growing class of automation tasks: support triage, moderation, intent classification, reranking, visual inspection, and guardrails for agents. In those environments, the best model is often not the most expressive one. It is the one that can classify fast, fail predictably, and slot cleanly into a workflow engine.
Liquid AI defines three question types in its technical materials: yes/no, multi-choice, and rubric-based scores. That makes the model closer to an inference component than a chatbot, and closer to an automation primitive than a general assistant. Teams already running workflow tools, internal queues, or edge devices should read this release through that lens.
Three numbers that explain where Open d1 fits
A short list captures why this launch stands out:
- 3.12B parameters: d1-3B is small enough to be practical, but large enough to post a 48.57 score on Decision Index v0.2.1, beating reported results from other sub-10B models in the source comparison.
- 8 ms per question on an RTX 4090: Liquid AI says d1-3B can answer one question in 8 ms on that hardware, or 16 ms without the cited compile optimisation.
- 587M parameters: d1-omni-600M extends the family into text-plus-image or text-plus-audio decisions, though Liquid AI did not publish latency figures for it at launch.
Those numbers suggest a broader market pattern. AI workflow automation is no longer only about replacing a person drafting text. Increasingly, it is about replacing brittle rules, regex layers, and multi-step classifier stacks with one decision pass.
That is especially relevant in customer support and operations. One model call can check refund eligibility, assign queue ownership, and score urgency on the same ticket. The source article notes that several questions can share one state in a single call, which reduces orchestration overhead.
Why zero-token models are attractive for production workflows
Most production failures in AI process automation do not come from model brilliance. They come from integration mess: inconsistent outputs, extra parsing logic, latency spikes, and edge cases that break downstream actions. A model that never writes prose removes part of that failure surface.
According to the Liquid AI launch post with links to the released models, the d1 family is built to read state once and return typed outputs. That matters in at least three settings:
- Customer support: triage, escalation, queue routing, and policy checks.
- Manufacturing and retail: image-based inspection, moderation, and shelf or camera review.
- Logistics and devices: voice-command routing and small-device intent detection.
This is also where AI automation agents may become more modular. Instead of letting a language model both reason and act, teams can split responsibilities. A generative model handles user-facing text; a decision model handles policy, routing, and guardrails. That architecture is often easier to test.
Liquid AI’s deployment story strengthens that case. The models are available via Transformers, have day-one llama.cpp support, and are positioned for the NVIDIA stack. That combination lowers friction for teams already standardised on server GPUs, workstation inference, or Jetson edge boards.
The latency story is strong, but the comparison needs caution
The most marketable number in the release is the 8 ms claim for one d1-3B question on an RTX 4090. Liquid AI also reports 16 ms on Jetson AGX Thor, 26 ms on AGX Orin, and 50 ms on Orin Nano. On Jetson AGX Thor, three questions over one state reportedly took 20 ms versus 16 ms for one.
Those are serious figures for AI task automation at the edge. They suggest that multimodal classification and inspection can move closer to cameras, terminals, and embedded systems without a cloud round trip.
But the comparison still needs discipline. The source itself notes that published latencies across open decision models use different hardware and workloads, so they are not directly comparable. The same caution applies to benchmarks. Liquid AI says it ran the official scorer for Decision Index itself rather than submitting to a public leaderboard. That does not invalidate the result, but it does mean buyers should treat it as a promising vendor-reported figure, not a final procurement answer.
This is where the market is likely to divide in 2027: not by who has the most agent demos, but by who can prove stable latency, reliable confidence calibration, and acceptable failure handling in a real queue or edge environment.
The biggest opportunity is not chat replacement
The easy mistake is to frame Open d1 as a smaller alternative to a chatbot. It is better understood as infrastructure for custom AI agents and operational workflows.
Consider the fit by use case:
| Use case | Why d1 looks useful | What still needs testing |
|---|---|---|
| Support triage | Multiple decisions over one ticket state, no text parsing | Confidence thresholds, routing accuracy by queue |
| Visual inspection | Reported 35 ms on a 384px image on Jetson AGX Thor in the source article | False positives, lighting variation, drift |
| Voice routing | 30-second audio support in d1-omni-600M | No published launch latency, English-only training scope |
| Agent guardrails | Typed yes/no or choice outputs fit policy checks | Coverage gaps, escalation paths |
This matters for AI analytics too. A typed output is easier to measure over time than free text. Operations teams can track class distribution, confidence changes, and failure buckets without another extraction layer. That makes continuous improvement simpler than in many chatbot-first deployments.
The non-obvious implication is organisational: zero-token models may reduce the amount of prompt engineering required for specific decisions, but they increase the importance of schema design. Teams must decide which questions exist, what labels are allowed, what score rubrics mean, and when confidence is good enough to trigger an action. The hard work moves upstream.
Where enterprises should be cautious
There are at least four real constraints.
First, d1-omni-600M is explicitly an early research release, and the absence of latency figures matters if the target use case is live audio routing.
Second, the Liquid AI launch post and linked license materials discuss LFM Open License v1.0 terms that allow free commercial use below $10 million in annual revenue. Larger companies need legal review before broad deployment.
Third, audio support is narrower than the headline suggests. Training covered English speaker-to-assistant requests only, and each request supports images or audio, not both together.
Fourth, this category only works when the business question is well-formed. If the task depends on open-ended judgment, negotiation, or long-form explanation, a decision model is the wrong tool.
For that reason, the best-fit internal service page here is AI Business Process Automation. It fits because Open d1 is most relevant when teams are implementing structured routing, triage, moderation, and inspection directly inside business workflows rather than experimenting with general-purpose chat experiences.
Conclusion
The trend behind Open d1 is clear: AI business automation is becoming more decision-centric, more multimodal, and more sensitive to latency at the workflow level. Liquid AI’s reported 48.57 benchmark score, 8 ms RTX 4090 latency, and immediate availability on mainstream tooling make this a release worth tracking.
The bigger signal is not that zero-token models will replace chat. It is that many production systems may work better when chat is only one layer, and the actual operational decisions are handled by smaller, typed, testable models.
Related reads
Martin Kuvandzhiev
Co-Founder & CEO, encorp.ai
CEO and Founder of Encorp.io with expertise in AI and business transformation
LinkedIn