ServicesSolutionsWorkProcessInsightsCompanyLife at SeikoContactFree AI audit

Home / Blog

Perplexity's Decisions API: Models That Output Probabilities, Not Text

Perplexity's Decisions API returns calibrated probabilities instead of text for classification and routing. What's real, the benchmark fine print, and the business case. Free AI audit.

Blue-lit network server racks in a dark data center aisle, representing AI routing and decision infrastructure

Image: Pexels / panumas nikhomkhai

On October 1, 2026, Perplexity opened a Decisions API powered by pplx-decider-v1-27b — a 27-billion-parameter model that, instead of generating text, returns calibrated probabilities over a fixed set of answers: yes or no, this team or that team, a ranked score. No parsing, no regex, no fragile JSON repair. (Perplexity Open-Sources 27B Decider, Edges Jev on 11-Test Panel — AI Weekly)

This column covered the category five days ago with TypeSafe's Jev. Perplexity's launch makes it a category — and the fine print is where the useful lessons are.

What the Decisions API actually is

A decision model is not a smaller chatbot. Perplexity's Hugging Face model card defines the interface: you send text (or text plus images) and a typed schema describing the questions to answer. Each question has a type — `noul` for yes/no, `choice` for multiple choice, `score` for ordered ranking (the three shapes are documented in Perplexity's API materials; the Hugging Face card shows `noul` and `choice` examples) — plus instructions and criteria. The model returns one probability per allowed option — no free text to parse. (perplexity-ai/pplx-decider-v1-27b · Hugging Face)

Perplexity's own example is a support-ticket router: the message "My Stripe integration keeps failing. Please help ASAP." goes in with criteria for `billing`, `technical_support`, and `sales`, and the model returns the owning team with a calibrated probability per option. A separate `noul` question measures urgency. That example names the target workload precisely: classification, routing, triage, and scoring at volume.

The model is a fine-tune of Qwen3.8-27B released as open weights under Apache 2.0, with a 262,144-token context window. The hosted API costs $0.04 per million input tokens with free output. Perplexity's docs show responses under two seconds for short prompts, stretching toward 23 seconds near the input ceiling — a quiet way of saying: keep your decision contexts small. (Perplexity Open-Sources 27B Decider, Edges Jev on 11-Test Panel — AI Weekly)

Three decider launches in one day

Perplexity did not launch into an empty field. On the same day, Cloudflare shipped Clef and Clef-flash, and AWS's Strands Labs released Strands Decider 2B — three launches aimed at the same classification-and-routing use cases Jev has been serving. (Perplexity Open-Sources 27B Decider, Edges Jev on 11-Test Panel — AI Weekly)

The practical signal: the category is filling in fast. There is now an open-weight 27B option, a 2B option from AWS's Strands Labs, and two models from Cloudflare, Clef and Clef-flash. Build your routing layer against a typed schema, not against one vendor's model.

The benchmark fine print

Perplexity reports pplx-decider-v1-27b at 85.71% overall on an 11-benchmark panel of 7,210 samples — edging out Jev at 84.51% and its own Qwen base at 74.76%. The widest gap is on RAGTruth, about using retrieved content faithfully: 88.80% against Jev's 77.27%. It also leads on FinancialPhraseBank, TabFact, and Circa. (perplexity-ai/pplx-decider-v1-27b · Hugging Face)

Read the fine print before quoting those numbers. They come from Perplexity's own panel, measured through the Perplexity API. The full table is published on the Hugging Face model card, but there is no independent rerun behind the headline figure — treat the numbers as the vendor's, not a third party's. Jev still leads on several of the eleven tests, including WinoGrande (90.70% vs 83.30%) and BBH (94.27% vs 82.80%). The honest reading: a strong decider model in the same tier as Jev, not a coronation. Both beat the Qwen base model by roughly 10 points on the same panel (85.71% and 84.51% vs. 74.76%) — a meaningful gap over the old pattern of prompting a chat model to "reply with JSON." (Perplexity Open-Sources 27B Decider, Edges Jev on 11-Test Panel — AI Weekly)

Decision model architecture: input state flows into a typed schema, returning calibrated probabilities per option for routing, scoring, and human review

What this means for your business

If your stack uses an LLM to "decide" anything — route a ticket, classify an email, score a lead — you likely have a brittle middle layer where a general model generates text and your code parses it. Every team that has shipped such a pipeline knows the failure list: the model explains instead of answering, the JSON breaks, a prompt tweak silently changes behavior elsewhere. The Decisions API attacks that layer directly: typed inputs in, calibrated probabilities out, no parsing step to maintain.

Three practical consequences:

1. Thresholds become a business decision, not an engineering hack. Calibrated probabilities let you write rules like "route automatically above 0.90, human review between 0.70 and 0.90, abstain below 0.70." That is how you put a decision model into production responsibly: define the abstention band, keep humans accountable for high-impact calls, and validate calibration per segment — not one global threshold and hope. And the threshold is a staffing model as much as a number: if 15–20% of your volume lands below 0.70, the human queue floods and your SLA dies. Price the abstention tier before you promise the automation. (Perplexity releases a low-cost model for probability-based decisions — Complete AI Training)

2. Classification gets much cheaper. At $0.04 per million input tokens with free output, a million short tickets costs on the order of ten dollars in model spend at typical prompt sizes — before you count the human review tier, which will dominate the budget. Compare that with sending every ticket to a full-size chat model and parsing the reply. Reserve expensive frontier models for steps that need real reasoning or generation; run the classification layer on a dedicated small model. Multimodal input extends the pattern in principle — the card accepts images alongside text — though Perplexity's published examples stick to text routing.

3. The self-host math is heavy. Apache 2.0 open weights are a real path to vendor independence, but the model card requires Python 3.12+ and a CUDA GPU with roughly 49 GiB just for weights plus working memory. Running a 27B prefill pass for a simple classification is expensive overkill. (Perplexity Decider 27B: Open weights decision model fine tune of Qwen3.8-27B — SignalHigh) For most teams the hosted API is the rational start — with the typed schema kept vendor-neutral, so you can swap in a smaller decider or your own fine-tune later. One caveat: schema portability is not calibration portability. Every swap means re-validating your thresholds against your own labeled data.

Two adoption notes an operations leader should hear. First, run it in shadow mode: keep your existing parse-based layer live, run the decider alongside it, and diff the disagreements before you cut over. Second, the hosted API means your tickets and emails — with their PII — leave your network. Answer the data-retention and processing questions before the engineering ones. And treat the criteria text as production logic: "Integration errors" and "Charges and refunds" are now config that routes real work, so version it and review changes like code.

One operational note: watch latency. Under two seconds for short prompts is fine for real-time routing; up to 23 seconds near the 262k-token ceiling is not. Decision contexts should be short by design — the relevant ticket, the relevant snippet, the defined options. If you are feeding whole conversation histories into a decider, you are using the wrong layer. (Perplexity Decider v1 27B: A 250K-Context Decision Model That Outputs Probabilities, Not Text — Medium)

The bigger story is the layer itself. Agents need exactly this underneath them: cheap, typed decision-making between the bigger models' reasoning steps — provided you log the inputs, schema versions, and probabilities so the decisions stay auditable. Three launches in one day say vendors are converging on the same pattern. If you are automating operations, standardize the decision layer first: it is where probabilities meet policy, and where the ROI arithmetic is easiest to check — once you've measured your error costs and your abstention rate.

Find out where AI can pay for itself in your operations — start with our free AI audit.

Sources

Want to know what this means for your stack? A free AI audit maps your workflows and shows where automation pays off — in your numbers, not ours.