OpenAI: GPT 4o Transcribe Diarize

Audio Transcription
openai/gpt-4o-transcribe-diarize

GPT-4o Transcribe Diarize is OpenAI's high-quality speech-to-text model, built on GPT-4o audio capabilities and released on 2025-10-15, with speaker diarization for multi-speaker recordings such as meetings and interviews. It accepts audio input and returns transcribed text, and is billed per token for input and output rather than per minute, which keeps costs transparent at the token level. Accessible via the OpenAI-compatible protocol through Ofox.

Context Window
128K
Max Output Tokens
128K
Released
2025-10-15
Capabilities
Audio Input
Available Providers
Azure
Supported Protocols
openai

Providers

Azure
Input Tokens
$2.5/M
Output Tokens
$10/M
Audio Input
$2.5/M
Web Search
$0.035/R
Protocols
openai/v1/audio/transcriptions

Code Examples

from openai import OpenAI
client = OpenAI(
base_url="https://api.ofox.io/v1",
api_key="YOUR_OFOX_API_KEY",
)
response = client.chat.completions.create(
model="openai/gpt-4o-transcribe-diarize",
messages=[
{"role": "user", "content": "Hello!"}
],
)
print(response.choices[0].message.content)

Frequently Asked Questions

OpenAI: GPT 4o Transcribe Diarize on Ofox.ai costs $2.5/M per million input tokens and $10/M per million output tokens. Pay-as-you-go, no monthly fees.