Lilith.
⌕
Editorial illustration: OpenAI narrowed GPT-6 Luna to three decision types, and a plugin brought them to the terminal
Lilith illustration · editorial remix

OpenAI has released a Decisions API powered by GPT-6 Luna, and Simon Willison has published version 0.1a0 of llm-openai-decisions for his LLM CLI. The API accepts text and images and returns three forms of structured output: the probability that a statement is true, a choice among predefined options or a score on a specified scale.

GPT-6 Luna returns a probability instead of a paragraph

Willison demonstrates an image query that returns a JSON object containing an answer type, an evaluation name and a probability. The plugin installs with llm install llm-openai-decisions and makes the model available in the same terminal workflow as other providers supported by LLM.

OpenAI charges $0.10 per million input tokens and nothing for output. The concept closely resembles TypeSafe AI's Jev, which also supports predicates, choices and scores. Willison's comparison prices Jev at $0.042 per million input tokens with text input, while GPT-6 Luna adds image input.

A dedicated decision layer simplifies routing and classification

A general generative model produces prose that an application often has to parse back into a category or number. The Decisions API gives developers a narrower contract. It fits request routing, queue prioritization, moderation, reranking and document assessment against fixed criteria.

The practical shift is not a smarter chat experience. It is an inexpensive component between unstructured input and deterministic application logic. Willison's plugin matters because it turns a product announcement into something developers can test repeatedly from the command line.

One decimal value can conceal both error and bias

Structured output makes integration easier while reducing visibility into a decision. The model returns a probability or score, not the reasons and specific signals that shaped it. In sensitive uses such as ranking job applicants or resolving customer disputes, a clean API should not be mistaken for auditability.

Low cost can also encourage broad deployment before a team measures false positives and false negatives on its own data. Evals need to cover boundary cases, language changes, image inputs and attempts to manipulate the model through submitted content.

Calibration on local data will determine practical value

Developers should test whether scores correspond to observed probabilities, how stable they remain under small prompt changes and whether safe thresholds can route uncertain cases to people. Jev comparisons should use the same task set because a lower token price says nothing about the cost of errors.

Image behavior is another important signal because OpenAI extends the format beyond Jev's text input. If GPT-6 Luna remains calibrated across both modalities and teams can continuously audit outcomes, narrow decision models may earn a distinct place alongside generative LLMs.

Lilith's verdict

GPT-6 Luna hands an application a clean number but leaves no witness statement beside it. Anyone using that number to sort people or disputes needs a human ready to challenge it.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗