A YouTube extension can now ask a model whether a thumbnail contains a wide-open mouth, then hide the video. A browser experiment asks the same API which action to take next. Both use OpenAI’s Decisions API to turn a judgment into something application code can act on.

OpenAI introduced Decisions at DevDay on 29 September and opened its public beta on 6 October. As of 11 October, that is twelve days since its introduction and five since the public beta began. There is already enough public work to examine how developers are using it.

This review draws on documentation, public code and authors’ experiments available on 11 October 2026. We have not reproduced the projects’ model evaluations.

How the API works

The Decisions API accepts text, images or both through POST /v1/decisions. Its current model is gpt-6-luna. A request supplies shared input and named questions, each with a defined answer type:

  • Predicate: the estimated probability that a condition is true.
  • Choice: one selection from options supplied by the developer.
  • Score: a probability-weighted average across ordered rubric levels, which can fall between levels.

The API can also refuse an individual question. Choice and score answers include distributions and a separate confidence field. These outputs need interpretation: application code decides whether to hide a thumbnail, route a ticket or ask a person to review it. Arbitrary JSON extraction and written explanations belong in the Responses API with Structured Outputs.

Readers of our Jev review will recognise the design: define the decision, give the model relevant evidence, then let software use the answer. The public projects below show several places to put that decision.

Six early uses

1. Filtering the YouTube feed

Ivan Campos’s YouTube Filters combines ordinary keyword rules with optional Decisions checks. Titles can be evaluated for categories such as clickbait or artificial urgency. Image questions handle visual features, including wide-open mouths and brush-style lettering on thumbnails.

The extension lets users hide a matching video or change its thumbnail’s appearance. Keyword matches avoid paid model calls, while previous AI results can be reused. Its request builder sends selected categories as named predicate questions, adding the thumbnail when an image check requires it.

This is a useful example of a model answering a small question inside an existing interface. The reader chooses what to filter and sets the threshold. A label such as clickbait remains dependent on those criteria; the more visual mouth or lettering checks give the project narrower questions to ask. The repository provides a manually installed Chrome extension, using the user’s own OpenAI key.

2. Choosing the next browser action

The decision-luna-browse-control experiment puts a decision model in a loop: inspect the current page, choose an action, execute it, then inspect again. Its test applications cover contacts, room reservations and support tickets in real Chromium.

The interesting part is how it checks completion. A workflow must survive saving, changing, closing and reopening a record, followed by an independent check of stored state. A convincing final screenshot alone does not pass.

In the author’s frozen controller comparison, Luna with DOM evidence, screenshots and bounded native waits completed nine of nine workflows: three repetitions in each local application. The repository reports a median successful completion time of 42.7 seconds and discloses an isolated rerun of the contact tests. The recorded results concern these synthetic applications and that controller version. They make the observation–action–verification loop worth studying, without establishing a success rate for unfamiliar commercial websites.

3. Reranking search results

Adrian Raudaschl’s RAG-Fusion experiment asks Decisions which retrieved documents help answer a query. Each request contains a pool of 50 candidate documents, with a relevance question for each. The author tries both yes/no judgments and a four-level relevance rubric.

On a sample of 200 NFCorpus queries, the rubric configuration reports 0.370 NDCG@10 for the single-query baseline. The hybrid-retrieval and query-fusion configuration adds 0.037, with a reported 95% paired-bootstrap interval of +0.014 to +0.062 for that improvement. NDCG@10 measures the quality of the first ten ranked results; it does not measure the accuracy of a generated answer.

The study also tests reversing the order, making it useful for examining where a decision model belongs in a retrieval pipeline. It covers one corpus and one Luna run per question type. Raudaschl discloses that he holds the RAG-Fusion patent, a relevant interest when reading his interpretation.

4. Routing work in n8n

The n8n Decisions community node brings the same idea into a visual workflow. Its Evaluate operation attaches answers to incoming items; Route sends each item to an output based on a decision. The package supports OpenAI alongside TypeSafe and OpenRouter.

Its documented support-ticket example chooses between billing, technical support and sales, with a separate path for uncertain answers. That gives a workflow builder a concrete starting point: keep the existing webhook and downstream integrations, then insert one classification step between them.

The adapter also handles a practical compatibility detail. TypeSafe-style yes/no questions become OpenAI predicates, and a refusal is preserved as a distinct result. This is a community integration with an example workflow; it is not evidence of a particular company’s support savings. Teams trying it should inspect both the model’s routing errors and the node’s configured fallback behaviour.

5. Sending pull requests to the right review

Vercel’s pull-request triage tutorial supplies a PR’s title, description, changed paths and patch text to Decisions. It asks which review track fits the change and whether callers may need migration guidance.

That is a more actionable question than asking whether a PR is “good.” A small change to a default value can require compatibility review even when the description calls it a cleanup. The example gives the model a needs_review option when the supplied evidence is insufficient.

Application code collects the diff and prints the returned decisions. The tutorial does not automatically approve or merge the PR. Its useful contribution is a bounded place for model judgment in an engineering process: deciding what a reviewer should examine next. Repository conventions and surrounding code still matter when a patch alone cannot establish the effect on callers.

6. Comparing forecasts with prediction markets

Model vs Market compares probabilities from OpenAI Decisions, Jev and Cloudflare Clef with Polymarket and Kalshi prices. Its description says the models receive recent news, with articles quoting betting odds removed before evaluation. A replay adds evidence article by article so readers can inspect how estimates change.

The OpenAI implementation asks a predicate question and reads the returned probability. Search is a separate component, supplied by Valyu. The decision endpoint does not gather the news itself.

The interface is interesting as a way to inspect sensitivity to evidence. Disagreement with a market price, however, is only a disagreement. To assess forecasting skill, we would want predictions recorded before events resolve, outcome scoring and calibration across many resolved questions. The project’s interface alone cannot establish a profitable trading strategy.

What the speed and price claims mean

OpenAI’s public-beta announcement claims decisions up to ten times faster than GPT-6 Luna through the Responses API. That comparison is with another OpenAI endpoint. It does not establish a tenfold improvement over Jev or over an entire browser or retrieval workflow.

The announced rate is $0.10 per million input tokens, with no output-token, cache-read or cache-write charges. Regional and long-context premiums can apply. Retrieval, screenshots, repeated calls and fallbacks still belong in the application’s cost calculation.

For a browser task, waiting for the page may dominate. For search, the size of the candidate pool changes how much evidence the model reads. A useful comparison measures the complete path to a successful result alongside the API call itself.

Choosing a first trial

Start with a decision your product already makes often enough to evaluate: selecting a review queue, filtering a thumbnail or scoring a retrieved passage. Keep the present system as the baseline and replay representative inputs, including ambiguous cases and known failures.

For routing, count expensive mistakes separately from harmless ones. For ranking, inspect what reaches the first results and what disappears. For browser actions, verify the saved application state after the workflow finishes. Record full workflow cost and latency, including retries and fallbacks; median and p95 timing will answer different operational questions.

A probability threshold should come from labelled examples and the consequences of a wrong answer. Give refusals and uncertain results an explicit destination. Initially, running decisions beside the existing workflow lets the team inspect disagreements before changing what users experience.

The first useful deliverable is a comparison the team can act on: which decisions improve, which inputs fail, and how much work remains around the model. These projects offer concrete starting points for that experiment.

← Back to blog