Summary

DeepSeek is transitioning V4 from preview to official status and introducing time-based API pricing: calls during Beijing-time peak windows cost roughly double the off-peak rate, while off-peak prices stay at today's flat baseline (V4-Pro at $0.435/M input and $0.87/M output). The legacy deepseek-chat and deepseek-reasoner endpoints retire on July 24, 2026.

What changed

DeepSeek confirmed V4's official launch with a new peak/off-peak pricing mechanism and set July 24, 2026 as the retirement date for its legacy chat and reasoner models, replacing the flat-rate pricing in place since the April 24 preview.

Why it matters

Peak/off-peak pricing is a notable shift for a lab that built its reputation on aggressive flat-rate economics, pushing cost-sensitive workloads toward off-hours batching and giving teams a lever to cut inference spend by scheduling. Retiring the legacy endpoints forces a migration deadline that API consumers must plan around.

Evidence excerpt

The API introduces a peak and off-peak pricing mechanism: during peak windows (9:00 a.m.-12:00 p.m. and 2:00 p.m.-6:00 p.m. Beijing time) prices double, while off-peak prices match today's baseline. Legacy deepseek-chat and deepseek-reasoner retire July 24, 2026.

Sources