# A model guide for the GPT-6 family

Date: 2 Oct 2026
Released by: OpenAI

## MediaRelease.co Summary

OpenAI published a practical guide on Oct. 2, 2026 for building with the GPT-6 family. It covers choosing among Astra, GPT-6.1 Sol and Luna, setting reasoning effort and speed modes, writing clear prompts and skills, steering long-running agent work, and using computer use in production workflows.

## Original release

Practical tips for getting the best results from GPT‑6 models while managing time and cost GPT‑6 is our most advanced suite of models yet, and offers you a choice of models for different kinds of work. Whether you’re turning an idea into a working prototype, building and testing a feature, or orchestrating multi-step workflows across code repositories, databases, and external APIs, this guide explains how to choose a GPT‑6 model, give it effective instructions, manage long-running work, and prepare for production. - Run effectively in production. Use caching and compaction to manage context and cost. Measure task success and latency, and plan for monitoring and data controls. - Match the model to your workload. Balance capability, cost, and latency by choosing the model, reasoning effort, and speed that fit the task. - Adjust your prompts and skills. Keep prompts, skills, and repository instructions consistent about what the model should deliver, what it can do independently, and what counts as done. - Keep long-running work on track. Use steering, async tools, and delegation to handle updates and independent work. Set clear boundaries for when the model should ask for input. 1. Run effectively in production Prepare your workflow for production Before deploying, there are several checks and best practices you’ll want to put into place. - Keep efficiency in mind. Cut context the task doesn’t need while keeping the evidence it does. Where your application supports it, run independent tasks together so one slow step doesn’t hold up unrelated work. - Reuse shared context through prompt caching for recurring work. Cached input tokens cost up to 95% less than uncached input tokens, depending on the model. Put stable instructions and reference material before changing task details, and keep tool definitions consistent. The caching dashboard and diagnostics guide help you see where that reuse breaks down. Include cache writes and any long-context rates when estimating the cost of a complete workflow. - For longer conversations, compaction reduces context size while preserving the state needed to continue. - Decide how you’ll monitor behavior and review the data controls for your application. - Test before deploying: Run representative tasks and measure task success, latency, and cost per successful task. Check out our API deployment checklist. Match the model to the workload Think of the model choice and reasoning level as an intelligence/ price tradeoff. - Model: - GPT‑6 Astra for the hardest reasoning work where maximum intelligence is needed. - GPT‑6.1 Sol for complex coding, research, and computer use. - GPT‑6 Luna for focused tasks at scale and everyday, repeated work with a clear goal, such as extracting invoice fields, classifying requests, or producing structured summaries. When evaluating the best model for the task, compare pricing for each model. - Reasoning level: In the API, choose how much effort the model spends on the task. - Low: Routine tasks, such as extracting facts or making small edits. - Medium: Work requiring judgment, such as planning a feature or comparing options. - High: Difficult debugging, deeper analysis, or careful review. - Extra high / Max: Test where supported when High falls short, and keep only if the improvement justifies the added time and cost. In the API, you can change reasoning effort mid-conversation without breaking cache. In Codex, start with the default reasoning level for that model, then lower it for simpler tasks or increase it for deeper analysis. - Speed: - In the API, use Fast mode when response time matters, such as in chat apps or coding tools. It provides faster, more consistent response times at a higher per-token cost than Standard processing. - In Codex and the API, use Ultrafast when faster responses are worth the premium, such as rapid coding iterations. It speeds up token generation independently of reasoning effort. Available for GPT‑6 Astra. 2. Adjust your prompts and skills Give the model a clear assignment “Models have gotten much better at understanding nuance and ambiguity, so overly specific guidance can now hinder results where it previously helped.” —Eric Provencher, Developer Experience at OpenAI Start with a clear assignment: the result you want, who it’s for, the relevant context and constraints, and what counts as done. Then review these four areas, summarized from Rethinking skills and prompts for GPT‑6 Astra , for a deeper dive into updating your instructions: - Create better skills: Keep descriptions short and explicit about when each skill should run, load supporting details only when needed, and replace rigid recipes with guidance suited to the models your team uses. - Update your AGENTS.md: Explain when particular documents and tests are relevant, and explicitly authorize safe routine workflows, such as running local tests with disposable data and no production access. - Set decision boundaries: State which actions can proceed independently and which require approval, replacing blanket “always ask” rules with clear boundaries. - Be prescriptive about persistence: Define what “done” includes—implementing the change, running it, inspecting the result, and fixing failures—and identify any decisions that require your review. For additional guidance, see reasoning best practices. Define the output you need Whether you’re working in Codex or building with the API, specify which decisions the model can make, when it should ask for input, and what a useful response looks like. Give the model enough direction to keep work moving without guessing at decisions that matter. Tell it which choices it can make and when to ask for your input, for example, it can choose how to organize a summary, but should check with you before changing the project’s scope. Describe what a useful response looks like, for example plain language, technical detail suited to your audience, and a short handoff covering what changed, what was checked, and what still needs attention. 3. Optimize long-running tasks Keep complex work moving With the GPT‑6 family of models, you can now take on tasks that span hours or days. Use the following features to better manage agents on long-running tasks. API In the API, use steering, asynchronous tools, and parallel work to keep long-running tasks moving. - Update instructions during a run: Mid-turn steering lets you send a correction through the Responses WebSocket API while the model works. Updates are queued; they don’t cancel running tools or undo completed actions. - Keep working while a tool runs: Asynchronous tool calling lets the model continue independent work while your app runs a slower task, such as tests. Your app returns the result when it’s ready. Wait for that result before starting work that depends on it. - Delegate independent subtasks: GPT‑6.1 Sol supports multi-agent workflows in the Responses API. It can assign independent work to subagents (such as investigating different parts of a codebase) and combine their findings into a final response. Multi-agent is currently in beta. Codex Long-running tasks can uncover decisions you wouldn’t necessarily anticipate in the initial prompt. Use clarification and steering to keep the work on course. - Answer questions as work progresses: With GPT‑6 Astra, Codex can ask for clarification while it works. Resolve questions that affect the next step, and specify which independent work can continue while you decide. If you’ll be away, tell Codex which tasks can continue and when it should pause for your answer. - Redirect work when requirements change: Steer the active task with new information, explaining what should change and what should stay the same. This helps avoid spending more time on an approach that no longer meets your needs. Leverage computer use to do more of the job Computer use lets GPT‑6 Astra, GPT‑6.1 Sol, and GPT‑6 Luna interact directly with websites and desktop apps, even applications without an API. For example, you can ask the model to investigate a bug, fix the code, and open your product in a browser to check that the fix works. Choose the simplest reliable way to do each step: - Use an API or connected tool when it can do the job directly. - Use computer use when the model needs to read a screen, click buttons, or fill in a form for you. If you’re building computer use into your own app, give the model a tool that can run code to control a browser or desktop. Playwright works with browsers; PyAutoGUI works with desktop apps. From testing to production: How teams are building with GPT‑6 Astra Harvey: more context, more useful drafts. Harvey combines court information, case law, firm documents, and a lawyer’s preferences to tailor its drafts. “We can give more context to the model and produce better and better structured outputs,” says cofounder Gabe Pereyra. Cognition: test results engineers can review. Cognition uses GPT‑6 Astra inside Devin to test software and return evidence. In an iPhone-game example, Devin produced a simulator recording and a report separating checks that passed from areas left untested, making the remaining work easier to see.

## Key details

- Issued by: OpenAI
- Published: 2 Oct 2026
- Publisher country: United States
- Subject country/region: United States
- Topics: AI
- Original source: https://openai.com/index/practical-guide-building-gpt-6
- MediaRelease.co URL: https://mediarelease.co/us/mr01185-a-model-guide-for-the-gpt-6-family-02102026.html

Original source: https://openai.com/index/practical-guide-building-gpt-6
MediaRelease.co canonical URL: https://mediarelease.co/us/mr01185-a-model-guide-for-the-gpt-6-family-02102026.html
