DeepSeek-V3 671B
DEEPSEEK:V3-671BMoE LLMOperationalLoad:68%SLA:99.99%Zero Queue
671B MoE (37B Activated) + Multi-Head Latent Attention
Base Rate
2 Credits / 1K Tokens
24h +38.90%
Groundbreaking 671B MoE architecture with 37B activated parameters, delivering world-class coding and reasoning at scale.
24h Total Calls
12.48Mcalls
+38.90% vs yesterday
Total Generations
0.01BTokens
+36.5% · Code & Reasoning 89%
Avg Latency (P50)
185ms
TTFT 110ms · 142 tok/s
Availability SLA
99.99%
Failover <45ms
Model Telemetry & Volume Chart
DEEPSEEK:V3-671BInvocations
546,026次+30.27%
Live Telemetry
Operational Performance
Production-grade Latency & Concurrency
P50 TTFT110 ms
P99 Latency380 ms
Throughput TPS
Peak Concurrency4,850 req
Packet Loss Rate0.0005%
Cost & Billing Matrix
Credits and Token Breakdown
Base Task Cost2 Credits
Input Cost / 1k Tokens¥0.001
Output Cost / 1k Tokens¥0.002
4K Upscale Multiplier1.0x
VIP Discount至高 40% 优惠
Technical Specifications
Parameters, Context & Resolutions
Architecture671B MoE (37B Activated) + Multi-Head Latent Attention
Parameters671 Billion Total
Context Window128,000 tokens
Max ResolutionN/A
Max DurationN/A
Global Edge Gateways
Live Region Latency & Load Balancing
East China (Shanghai)
12ms
South China (Guangzhou)
16ms
Hong Kong Gateway
28ms
Asia Pacific (Singapore)
48ms
US West (Silicon Valley)
135ms
EU Central (Frankfurt)
154ms
