DeepSeek V4.1 Flash
Chatdeepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash is DeepSeek's latest efficiency-optimized Mixture-of-Experts model with a 1M-token context window and up to 384K output tokens. It adds native multimodal vision understanding, supports thinking mode (on by default), tool calls, JSON output and prompt caching, and per DeepSeek surpasses V4 Pro on quality, cost and speed. Served upstream under the official model id deepseek-flash.
Context Window
1M
Max Output Tokens
384K
Released
2026-09-10
Capabilities
VisionFunction CallingReasoningPrompt Caching
Available Providers
DeepSeek
Supported Protocols
openaianthropic
Providers
DeepSeek
Input Tokens
$0.3/M
Output Tokens
$1.2/M
Cache Read
$0.006/M
Protocols
openai
/v1/chat/completions/v1/responsesanthropic
Code Examples
from openai import OpenAIclient = OpenAI(base_url="https://api.ofox.io/v1",api_key="YOUR_OFOX_API_KEY",)response = client.chat.completions.create(model="deepseek/deepseek-v4.1-flash",messages=[{"role": "user", "content": "Hello!"}],)print(response.choices[0].message.content)
Related Models
Frequently Asked Questions
DeepSeek V4.1 Flash on Ofox.ai costs $0.3/M per million input tokens and $1.2/M per million output tokens. Pay-as-you-go, no monthly fees.
DeepSeek V4.1 Flash supports a context window of 1M tokens with max output of 384K tokens, allowing you to process large documents and maintain long conversations.
Simply set your base URL to https://api.ofox.io/v1 and use your Ofox API key. The API is OpenAI-compatible — just change the base URL and API key in your existing code.
DeepSeek V4.1 Flash supports the following capabilities: Vision, Function Calling, Reasoning, Prompt Caching. Access all features through the Ofox.ai unified API.