Primary sourceAIArticle··5 min read

GPT-6: OpenAI's Playbook for Moving From Prototype to Production

Three models, five reasoning levels, two speeds and agents that work for hours: OpenAI publishes the manual for its new family.

GPT-6: OpenAI's Playbook for Moving From Prototype to Production
Source : OpenAI · OpenAIView original ↗

In brief

OpenAI has released a practical guide for startups building with the GPT-6 family: GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna. The document explains how to pick the right model and reasoning level, rewrite your prompts and skills, steer long-running tasks, and control costs in production. The takeaway is clear: the developer becomes the project manager of autonomous agents.

🍺 Bar-stool version

OpenAI drops three models at once and, knowing full well you're going to grab the priciest one just to sort out your invoices, publishes a manual to tell you which one to actually pick. The underlying message is simple: talk to the AI like a brilliant intern, tell it what you want, what it's allowed to do on its own, and when the job is actually done. On top of that, these agents can now grind away for hours, even days, and even click buttons for you, making them the only coworkers who never ask for a coffee break. What really matters is that the real skill isn't writing code anymore, it's writing a spec that nobody used to bother reading.

Key takeaways

  1. 1

    The GPT-6 family comes in three models: Astra for the toughest reasoning, GPT-6.1 Sol for complex code, research and computer use, Luna for high-volume repetitive tasks.

  2. 2

    Reasoning effort is set across five levels (Low, Medium, High, Extra high / Max), to be treated as an intelligence/price tradeoff.

  3. 3

    Two paid speed modes: Fast for quick, stable response times, Ultrafast (Astra only) which speeds up token generation independently of reasoning.

  4. 4

    Prompt caching cuts the cost of cached input tokens by up to 95%, provided stable instructions are placed before changing details.

  5. 5

    OpenAI recommends replacing "always ask" rules with explicit decision boundaries and precisely defining what "done" means.

  6. 6

    For long-running tasks: in-flight steering via the Responses WebSocket API, asynchronous tool calls, and multi-agent workflows (in beta on GPT-6.1 Sol).

  7. 7

    All three models can operate websites and desktop apps via computer use, including software with no API.

A lineup built as a tradeoff

The guide sets the frame right away: choosing a GPT-6 model means trading off capability, cost, and latency. GPT-6 Astra is reserved for the hardest reasoning problems, where maximum intelligence is required.

GPT-6.1 Sol targets complex code, research, and computer use. GPT-6 Luna is aimed at clear-purpose bulk work: extracting fields from invoices, classifying requests, producing structured summaries.

Reasoning level is set separately. Low for extraction or small edits, Medium for planning a feature, High for hard debugging. Extra high / Max should only be kept if the gain justifies the extra time and cost.

On speed, Fast mode offers quicker and more consistent responses at a higher per-token rate. Ultrafast, available only for Astra in Codex and the API, speeds up generation regardless of reasoning effort.

Production is first and foremost about context and costs

OpenAI emphasizes efficiency: trim unnecessary context while keeping the evidence that matters, and run independent tasks in parallel so a slow step doesn't block the rest.

Prompt caching is highlighted: cached input tokens cost up to 95% less depending on the model. The recipe is to place stable instructions and reference documents before variable details, and to keep tool definitions constant.

For long conversations, compaction reduces context size while preserving necessary state. Before any deployment, the guide recommends measuring success rate, latency, and cost per successful task, and planning for monitoring and data controls.

Prompts, skills and AGENTS.md: writing an actual spec

The core advice, attributed to Eric Provencher (Developer Experience at OpenAI), fits in one sentence: give the model a clear assignment. The expected result, its recipient, the context, the constraints, and what counts as done.

Skills should have short descriptions specifying when to trigger them, loading details only when needed, and trading rigid recipes for guidance adapted to the models being used.

The AGENTS.md file should explain when a given document or test is relevant and explicitly authorize safe routines, such as running local tests on disposable data without production access.

Finally, the guide asks to be prescriptive about persistence: "done" includes implementing, running, inspecting the result, and fixing failures. The model can, for example, choose how to organize a summary, but must check in before changing the project's scope.

Agents that work for hours, even days

OpenAI states that with GPT-6, you can hand off tasks spanning hours or days. In the API, in-turn steering lets you send a correction via the Responses WebSocket API while the model is working; updates are queued without canceling in-flight tools or actions already taken.

Asynchronous tool calls let the model keep making progress on independent work while a slow task, like a test suite, runs on the application side. GPT-6.1 Sol also supports multi-agent workflows, still in beta: it delegates to sub-agents then merges their findings.

In Codex, Astra can ask clarifying questions along the way. The guide advises specifying which tasks can continue while you think, and explicitly redirecting work when requirements change.

Computer use, as a last resort

All three models can interact directly with websites and desktop applications, including ones without an API. Example given: investigate a bug, fix the code, then open the product in a browser to verify the fix.

The proposed rule is pragmatic: go through an API or a connected tool whenever possible, and reserve computer use for cases where you need to read a screen, click, or fill out a form. To integrate it into your own application, OpenAI cites Playwright for the browser and PyAutoGUI for desktop.

“Give the model a clear assignment.”
“Think of the model choice and reasoning level as an intelligence/ price tradeoff.”
“Cached input tokens cost up to 95% less than uncached input tokens, depending on the model.”

Why it matters

This guide reveals more about OpenAI's strategy than a simple product sheet would. The proliferation of dials (model, reasoning effort, Fast, Ultrafast, caching) turns using an LLM into an economic optimization exercise, and the vendor is openly betting that cost per successful task is the metric that matters. The second message is cultural: the developer's job is shifting toward delegation, with decision boundaries, definitions of "done," and agents you redirect mid-task. It's coherent, but it deserves a critical eye. A guide written by the vendor naturally pushes toward its own tools (Responses API, Codex, premium modes) and says nothing about the real limits of these agents on multi-day tasks. Multi-agent is still in beta, and the promise of sustained autonomous work remains to be verified on real-world cases, along with the corresponding bill.

Free account

You just read an AI Sources article

Create a free account: a month of archives in full, your own sources summed up like this one, your notes and highlights.

#openai#gpt-6#agents#codex#api#llm
Original source
A model guide for the GPT-6 family
OpenAI
Open the article ↗

For you

Put it to work on your sources.

Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next