A new shape of model, and why marketers should care
Most of the AI you have plugged into your marketing stack so far speaks in sentences. You send it a prompt, it sends back text, and somewhere downstream you or a script parse that text into a decision: pause this ad group, add this user to that segment, raise this bid. The text is the awkward middleman. It is expensive to generate, slow to stream, and it hallucinates in ways that only surface when your parser chokes on the 3am run.
Last month a model called Jev, from TypeSafe AI, put a different shape on the table. As Simon Willison described it, Jev still accepts text inputs, but instead of text output it returns floating point numbers corresponding to categories, yes/no questions, ratings, and associated confidence scores. TypeSafe frames it as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. Willison, following Maggie Appleton, prefers the name "decision models" over TypeSafe's own "System One models." That is the framing we find most useful for marketing work, so it is the one we will use here.
What a decision model actually does
The mechanics are simple enough to hold in your head. You compose a "state" object, a string, an array of strings, or a set of name-value pairs, describing an article, a customer, or any other record. You send it to the API with one or more questions and get an answer per question. Willison lays out three question types: Yes/No questions (Jev calls them "Noul" questions, short for Bernoulli), which return a number between 0 and 1 for how confident the model is that a statement is true; Choice questions, which pick one option from a set and return a confidence score plus a probability distribution; and Score questions, which place a value along a numeric range you define.
Questions are evaluated in parallel, so asking fifty of them takes about as long as asking one. And it is cheap. Willison reports that Jev charges only for input, with output free, and quotes the input price of their first model at $0.042 per million tokens, which he notes is cheaper even than OpenAI's GPT-5 Nano at $0.05 per million. When a judgment costs that little, the economics of running a judgment on every row of your data change completely.
Where this maps onto marketing decisions
Think about how much of daily campaign work is really classification and scoring dressed up as text generation. Willison points to spam detection, suggesting labels, prioritization and ranking as natural fits, and to search reranking as one he has been experimenting with: fetch a hundred candidates with a cheap algorithm, then score them for relevance. Swap the nouns and you have a marketing backlog.
- Audience segmentation. "Does this account match our ICP?" is a Noul question. "Which lifecycle stage is this contact in?" is a Choice question. Instead of a brittle rules engine, you get a calibrated probability per record.
- Ad decisioning. "Is this search term relevant to what we sell?" run across every query in a report is a negative-keyword workflow that used to eat an afternoon.
- Lead and creative scoring. Score questions with your own described levels turn "rate this lead 1 to 5" into something you can run on the whole table for a few cents.
This is the deterministic-agent territory we keep coming back to. A text model that occasionally invents a bid is a liability inside an automated loop. A decision model that returns a confidence score is something you can actually build a gate around: act above 0.9, queue for review between 0.6 and 0.9, ignore below. We wrote about this shift in how to use AI to transform marketing workflows, not just speed them up, and decision models are the clearest example yet of moving work into the system rather than parking a chatbot beside it.
The black box gets darker
Willison is candid about the cost, and so should we be. Jev represents a regression further towards black box machine learning. Regular LLMs are already black boxes, he notes, but at least you can ask them to justify a decision, even if the justification is not reliable. Jev does not give you that. Put in all the text you want, and the only thing coming back is a floating point number. If Jev marks something as spam, you do not learn which signals tipped it off.
For marketing that matters twice over. First, bias. Willison is explicit that concerns about bias should be front and center, and that he hopes nobody uses this class of model to rank job applicants, because that single number can conceal unseen bias. Anything where you are scoring people, and audience segmentation is exactly that, inherits the warning. Second, accountability. When a segment or a bid is set by a number with no rationale attached, your only defence is measurement. Willison's own conclusion is the one to internalise: evals and structured experiments matter even more for decision models than for regular LLM projects, and because the model is cheap, running hundreds or thousands of test prompts costs just a few cents. Cheap decisions are only safe if you spend some of the savings on watching them.
You still need the pipes
Here is the part nobody's launch video mentions. A decision model is a scoring function. It does not know your conversion data, it does not close the loop, and it does not tell you whether the segment it drew actually converted better. That is your instrumentation's job. The confidence score from a Noul question is only worth acting on if you can trace, weeks later, whether the accounts you gated in outperformed the ones you gated out.
So the workflow we would actually build looks less like "replace the LLM with Jev" and more like a chain: pull the records from BigQuery, shape them into state objects, send batched questions, write the scores and confidences back to a table, act on the ones above threshold, and measure outcomes against the ones you held back. Orchestration tools like n8n make the plumbing between those steps mundane, which is the point. If you want a sense of how much of the agency operating model this pattern is quietly rewriting, we covered that in what eight agencies are doing differently after rebuilding around AI.
The honest takeaway
Decision models are not a smarter chatbot. They are a cheap, fast, calibrated classifier with a natural-language interface, and that is genuinely useful for the large chunk of marketing automation that is really just repeated judgment at scale. The catch is that they hand you a number and nothing else, which raises the bar on both bias vigilance and measurement. The teams that win with this shape will be the ones who already have the analytics discipline to check whether the numbers were right.
If you are weighing where decision models or MCP-connected agents fit into your stack, and want the measurement layer built so you can actually trust the decisions, talk to Bitegrico about your analytics and automation implementation. That is the work we do.