GPT-Live API: Release Timeline, Endpoints, and Developer Access
A current implementation guide to the GPT-Live API: supported models, live endpoints, exact-second voice billing, connection options, delegation, and the evidence boundaries that still matter when you plan a production integration.
For builders, the relevant developer surface is now documented as the GPT-Live API. The official model entry lists gpt-live-1 on v1/live/sessions, with audio and text in, audio and text out, and full-duplex conversation. The GPT-Live guide turns that surface into connection and session instructions; consumer ChatGPT access remains separate from API billing and project eligibility.
What We Know About the GPT-Live API So Far
It's a Real-Time Streaming Endpoint
The official live model entry lists v1/live/sessions as the supported endpoint for gpt-live-1. The live guide describes full-duplex conversation, so the model can listen and speak at the same time while your application receives live events for the conversation and delegated work.
Backed by the Same Model Family as the ChatGPT App
Use the API model ID gpt-live-1 for the documented live voice session; see the official gpt-live-1 model page. The official model pages used here do not establish a gpt-live-1-mini API model, so this page does not present that name as an API option. The separate gpt-live-transcribe model is for low-latency speech-to-text through v1/realtime/transcription_sessions.
Availability Is Staged
The current OpenAI Developer documentation describes a usable API surface for projects with access, while rate limits and account eligibility still depend on the project. The model page lists the Free API tier as unsupported for gpt-live-1; check the current page before relying on a specific limit.
GPT-Live and the Realtime API remain separate documented voice surfaces. The GPT-Live migration guide explains how to move from Realtime to Live, while the Realtime guide remains available for Realtime integrations. Choose the documented surface that matches your model, transport, and existing architecture.
Expected Capabilities for the GPT-Live API
Bidirectional Audio Streaming
GPT-Live-1 accepts audio and text input and returns audio and text output; its model entry does not list image or video input/output. You can connect a browser with WebRTC, a server with WebSockets, or a phone workflow with SIP. The connection-specific guides define the event and media details.
Function Calling and Tools Mid-Conversation
The model page lists streaming and function calling as supported. Keep permission checks, confirmations, and business actions in your application. The live guide describes an optional server-side sideband connection that can observe or control an existing session while the primary connection continues to carry the conversation.
Handoff to Other Models
Delegation separates the voice model from backend reasoning and tools. With Responses delegation, GPT-Live calls the Responses model you select. With Client delegation, your application supplies the context and runs any model, agent harness, or service it operates, then sends the result back to GPT-Live.
Multilingual ASR + TTS in One Pipeline
For the documented live voice model, treat audio conversation and transcription as separate API surfaces. gpt-live-1 is the full-duplex voice model; gpt-live-transcribe returns text transcript deltas through v1/realtime/transcription_sessions. This page does not claim a fixed language-quality tier or built-in translation behavior without a source for that exact claim.
- WebRTC: browser voice applications with microphone and speaker media tracks plus a data channel for events.
- WebSockets: server-side audio and control events on a trusted connection.
- SIP: telephony connections, with optional sideband control from your backend.
- Sideband: an additional server connection described in the live guide for monitoring or commands; audio remains on the primary session connection.
Pricing Signals and Cost Considerations
The official model page and pricing page list gpt-live-1 voice sessions at $0.05 per minute, billed by the second without rounding up to a whole minute. The separate gpt-live-transcribe model is listed at $0.017 per minute for realtime transcription.
- Voice duration cost:
session seconds × 0.05 / 60. - Backend Responses model usage and tool usage are separate charges under the official pricing guidance.
- A one-second session is billable even when a two-decimal display rounds the dollar amount to $0.00; use sufficient precision for reconciliation.
Do not turn the voice-session rate into a monthly subscription price. If you are planning a month, enter the expected monthly session seconds and label the result as a monthly projection. See the pricing page for the separate consumer-plan discussion.
How to Get Developer Access to the GPT-Live API
The current implementation path is to use the official GPT-Live guide, confirm that your project can use the listed model, keep the API key on a trusted server, and start with the connection guide that matches your application. The newsletter form at the end of this page is only for independent GPT-Live Hub updates; it is not an OpenAI access request or application.
- Open the official gpt-live-1 model page and confirm the current endpoint, rate limits, and supported tier for your project.
- Choose WebRTC, WebSockets, or SIP for the primary connection.
- Keep project credentials on a trusted server and test session close, permissions, tool confirmations, and usage reporting.
- For backend work, choose Responses delegation or Client delegation in the delegation guide.
Microsoft Foundry also has a directly verified GPT-Live WebSocket article that documents an Azure resource, a gpt-live-1 deployment, and Azure live-session configuration. This page does not infer regional availability or concurrency limits from that article.
The official guide set also includes Getting started, Prompting, Managing sessions, Delegation and tools, Migrate, and Partner integrations.
Build Ideas Worth Lining Up First
Voice-first support agents
Always-on customer support that can answer, look up records, and escalate, all in one voice session.
Real-time interpreters
Travel, healthcare, and field-service translation between two languages.
Hands-free coding copilots
Voice-driven code review and shell operations for low-vision or accessibility-first users.
Accessibility tooling
Low-vision or motor-impaired users get a full conversational interface by voice.
Voice companions for kids and language learners
Patient, pause-aware conversation partners, subject to appropriate product safeguards.
Risks and Things to Watch
Latency on Real Hardware
Latency depends on the connection, audio pipeline, backend work, and workload. Test on the devices and networks you support, and provide a clear recovery path when a session or delegated task fails.
Policy and Safety
Do not assume consumer ChatGPT behavior or an undocumented safety tier maps directly to the API. Apply the current API terms and safety guidance to your product, validate tool permissions, and handle refusals or interruptions without presenting them as successful actions.
Cost at Scale
Track live session seconds, backend model usage, and tool charges as separate line items. A persistent voice connection can create a meaningful duration bill even when the visible voice output is short, so close sessions deliberately and reconcile final usage.
FAQ on the GPT-Live API
The official model page describes gpt-live-1 and the v1/live/sessions endpoint. Confirm project access, current rate limits, and supported tiers before shipping.
gpt-live-1 voice sessions are listed at $0.05 per minute, billed by the second without rounding up. Backend model and tool usage is separate. gpt-live-transcribe is listed at $0.017 per minute.
Commercial use depends on the current OpenAI terms, account access, and your implementation. Review the applicable terms and keep API credentials on trusted servers.
Final Notes on the GPT-Live API
For developers, the documented GPT-Live surface is a full-duplex voice session connected to a backend workflow you control or delegate. Start with the official model and connection guides, measure voice seconds separately from backend and tool usage, and keep consumer ChatGPT claims out of API implementation decisions. See pricing → Back to GPT-Live guide →
Sources
Sources below link directly to the official documentation used for the implementation details. OpenAI API pages were rechecked on 2026-09-20; Microsoft Learn is cited only for the bounded Azure connection statement.
- OpenAI Developer Docs — GPT-Live 1 model, endpoint, rate limits, and pricing
- OpenAI Developer Docs — GPT-Live-Transcribe model and transcription endpoint
- OpenAI Developer Docs — API pricing
- OpenAI Developer Docs — Getting started with GPT-Live
- OpenAI Developer Docs — Prompting
- OpenAI Developer Docs — Managing sessions
- OpenAI Developer Docs — Delegation and tools
- OpenAI Developer Docs — Migrate from Realtime to Live
- OpenAI Developer Docs — Partner integrations
- OpenAI Developer Docs — Realtime API
- OpenAI Developer Docs — WebRTC; WebSockets; SIP
- Microsoft Learn — Use GPT-Live for real-time voice (Azure connection evidence; regional availability and concurrency not inferred)