Market Intelligence Report

Kimi vs DeepSeek

Kimi K3 vs DeepSeek V4: official API pricing, MIT vs open-frontier weights, multimodal agents, Reddit/HN sentiment. When to pick Flash, Pro, or K3.

The Contender

Kimi

Best for AI Models

Starting Price Contact
Pricing Model freemium
Kimi

The Challenger

DeepSeek

Best for AI Writing

Starting Price Contact
Pricing Model freemium
Try DeepSeek

The Quick Verdict

Kimi K3 wins native vision, long-horizon agents, and open-frontier intelligence at $3/$15 miss/out ($0.30 cache-hit). DeepSeek V4 Flash wins pure text cost ($0.14/$0.28 per 1M miss/out) with MIT open weights today.

Independent Analysis

Feature Parity Matrix

Feature Kimi DeepSeek
Pricing model freemium freemium
deepseek v3 model Yes
cost effectiveness High (fraction of the cost)
r1 reasoning model Yes
multilingual support Yes
performance rivals gpt4 Yes
open source availability Yes
code generation capabilities Yes
Quick Answer

DeepSeek V4 Flash wins pure text cost ($0.14/$0.28 per 1M miss/out) with MIT open weights today. Kimi K3 wins native vision, long-horizon agents, and open-frontier intelligence at $3/$15 miss/out ($0.30 cache-hit). Route both: Flash for volume, K3 or V4 Pro for hard sessions.

Quick verdict

Kimi K3 (Moonshot AI) and DeepSeek V4 are the two open-weight Chinese labs Western builders actually put in production in mid-2026. This page is K3 specifically versus DeepSeek’s V4 line—not a vague “Kimi brand vs DeepSeek brand” blur. K3 is a 2.8T-parameter multimodal MoE with 1M context and frontier-class API pricing; DeepSeek V4 is a pair of MIT-licensed text models (Flash + Pro) engineered for cheap million-token serving.

Pick DeepSeek when pure text volume and permissive open weights dominate: V4 Flash is the default cheap implementor ($0.14 miss / $0.28 out per 1M; cache-hit input at $0.0028), V4 Pro is the reasoning/coding step-up ($0.435 / $0.87) with the same 1M context and near-free cache hits. Pick Kimi K3 when you need native vision, long-horizon agent sessions, Kimi Code / Work product paths, or the current open-frontier ceiling—and you will pay roughly Claude-Sonnet-class rates ($3 miss / $15 out, $0.30 cache-hit) for it.

One-liner

DeepSeek V4 wins the text-only cost war. Kimi K3 wins multimodal open-frontier sessions and productized agents. Most serious stacks route both by task.

Side-by-side

DimensionKimi K3 (Moonshot)DeepSeek V4
CompanyMoonshot AI (Beijing)DeepSeek (Hangzhou)
Flagship SKU(s)kimi-k3 (single max-thinking default at launch)deepseek-v4-flash + deepseek-v4-pro
Scale2.8T total MoE; 16 of 896 experts active (vendor)Pro ~1.6T/49B active; Flash ~284B/13B active
ModalityNative vision / multimodalText-only on first-party V4 API
Context1,048,576 tokens1M both SKUs; max output 384K
Open weightsPromised full dump by 2026-07-27; K2-line Modified MITLive MIT on Hugging Face now
API floor (1M tok)$0.30 cache-hit / $3 miss / $15 outFlash $0.0028/$0.14/$0.28; Pro $0.003625/$0.435/$0.87
Concurrency (official)Not published as a single global number; product + API quotas applyFlash 2500 / Pro 500
Product shellkimi.com, Kimi Code, Kimi Work, apps, EnterpriseThin chat + API platform; harness is your product
Best default forHard agents, vision-in-loop coding, CN/bilingual knowledge workCheapest serious text/coding at scale

What each product is in 2026

Kimi K3 is Moonshot’s open-frontier flagship, introduced mid-July 2026 as a 2.8-trillion-parameter MoE built on Kimi Delta Attention (KDA) and Attention Residuals, with Stable LatentMoE (16 of 896 experts), native vision, and a 1M-token context window. Moonshot positions it as the first open 3T-class model, still trailing Claude Fable 5 and GPT 5.6 Sol overall while matching or beating other open and mid-closed peers on many agentic/coding suites. Surfaces: kimi.com, mobile apps, Kimi Code (select K3 via /model), Kimi Work desktop, and OpenAI-compatible API on platform.kimi.ai with model id kimi-k3. At launch, max thinking effort is the default; lower-effort modes were scheduled as follow-ups. Full public weights were scheduled for July 27, 2026 with vLLM KDA/prefill-cache work—until the HF repo is live, treat “open” as a roadmap commitment plus API access, not a self-host guarantee.

DeepSeek V4 is the cost-and-open-weights specialist that followed the R1/V3 shock. Two production SKUs: V4 Flash (smaller MoE, high concurrency, rock-bottom price) and V4 Pro (larger MoE, deeper reasoning/coding). Official docs list 1M context, up to 384K max output, thinking and non-thinking modes, tool calls, JSON mode, FIM (non-thinking), and OpenAI- plus Anthropic-compatible base URLs. Weights are already on Hugging Face under plain MIT (~1.6T/49B active Pro; ~284B/13B active Flash class per report/model cards). The product UI is intentionally thin—free chat plus a billing platform—so most production “DeepSeek experience” is whatever harness you wire (OpenCode, Hermes, Claude Code proxies, OpenRouter, self-host).

Watch out: Comparing “Kimi K3” to “DeepSeek” without naming Flash vs Pro is useless. K3 is not K2.6; Flash is not Pro. Price and quality answers flip by SKU—budget the model id, not the brand.

Pricing and real cost (TCO)

Sticker math favors DeepSeek hard on raw text. K3 is priced like a frontier closed model, with a large cache-hit discount that only helps if your prefixes are stable. Numbers below come from official pages as of research—re-check before you lock a budget.

DeepSeek V4 API (official)

  • deepseek-v4-flash — cache-hit input $0.0028, cache-miss $0.14, output $0.28 per 1M tokens; 1M context; concurrency limit 2500.
  • deepseek-v4-pro — cache-hit $0.003625, miss $0.435, output $0.87 per 1M; 1M context; concurrency 500. Earlier temporary cuts were reported as permanent on the Pro SKU in press coverage—still verify the live table.
  • Legacy names deepseek-chat / deepseek-reasoner map to Flash non-thinking / thinking and deprecate 2026-07-24 15:59 UTC—migrate IDs.
  • Hosted chat remains free with product limits; production cost is almost entirely API top-up.
  • Community threads discuss peak/off-peak multipliers (Beijing business hours). Treat that as operational risk: re-check official policy before multi-region capacity planning.

Kimi K3 API + membership

  • kimi-k3 API (official blog + pricing docs) — cache-hit input $0.30, cache-miss $3.00, output $15.00 per 1M tokens; context 1,048,576. Moonshot markets >90% cache-hit rates on coding workloads via Mooncake disaggregated inference.
  • OpenRouter lists the same $3 / $15 miss/out with $0.30 cached input for moonshotai/kimi-k3—useful for failover, not a free lunch.
  • Membership (Adagio free → higher Allegro/Vivace tiers) still gates consumer agent/Code/Claw quotas separately from pure API metering. Heavy always-on agents can hit membership limits even when API math looks fine.
  • For context, prior flagship K2.6 sat around $0.16 cache-hit / $0.95 miss / $4 out with ~262K context—K3 is a clear step up in both capability claims and sticker price.

Order-of-magnitude: 10M output tokens on V4 Flash is ~$2.80; on V4 Pro ~$8.70; on Kimi K3 ~$150—before cache games. That is why Flash owns volume pipelines and K3 owns sessions you actually supervise.

TCO notes: Cache-hit rates rewrite both bills. DeepSeek’s sub-cent cache hits make multi-turn agent loops extremely cheap when the prompt prefix is stable (HN reports of multi-million-token burns for under a dollar). K3’s $0.30 hit vs $3 miss only pays if coding/agent harnesses keep long system prompts warm; Moonshot’s “>90% hit on coding” claim is workload-specific. Independent and Reddit notes also stress that token count per task matters: a model that is cheaper per token but more verbose can lose on cost-per-solve; AA-style token-per-task charts put K3 in a more efficient band than some closed peers while still using more tokens than the most terse frontier models.

Features that actually differ

Open weights & license. DeepSeek V4 Pro/Flash weights are public under MIT today—cleanest commercial redistribution story among the two for self-host and fine-tune. Kimi’s prior open models ship Modified MIT (branding obligations matter mainly at huge commercial wrapper scale). K3’s full dump was scheduled for July 27, 2026; until then, API is the primary access path. Local full-size K3 is a multi-TB quant problem even for well-equipped homelabs.

Modality. K3 ships native vision and vision-in-the-loop coding (screenshots, game/frontend, CAD-style loops in Moonshot case studies). DeepSeek V4 first-party is text-only—fine for pure code/docs, a hard stop if the agent must “see” UI state without a second vision model.

Architecture (directional). K3: KDA + AttnRes + ultra-sparse MoE + MXFP4 QAT story for serve efficiency at 2.8T. DeepSeek V4: hybrid Compressed Sparse Attention / Heavily Compressed Attention, Manifold-Constrained Hyper-Connections, DeepSeekMoE; report claims ~27% of V3.2 single-token FLOPs and ~10% KV cache at 1M context for Pro-class efficiency. You do not need the papers to ship—but they explain why both labs can advertise 1M context at all.

Agents & tools. K3’s product bet is long-horizon coding, multi-agent knowledge work (Kimi Work widgets/dashboards), and Code/Claw-style sustained tool use; OpenClaw docs already default new users toward moonshot/kimi-k3. DeepSeek ships solid thinking-mode tool calls and lets the ecosystem harness do the rest—OpenCode users often call Flash “magical” for implementor loops while routing planning to Pro or a closed model.

Coding quality (directional, harness-sensitive). Moonshot reports strong agentic coding numbers (e.g. DeepSWE ~67 with common harness notes, Terminal-Bench 2.1 high 80s under KimiCode)—but independent writeups stress harness confounds and missing common-board entries. Artificial Analysis places K3 high on intelligence (~57 index; joint top-tier coding-agent index entries around 57). DeepSeek community consensus: Flash for volume implementor work; Pro when the task is actually hard; some coders underwhelmed relative to closed Sonnet/Opus or Kimi on long plans. Run your repo.

Language & UX. Practical writeups still give Moonshot products an edge for idiomatic Chinese and bilingual product UX; DeepSeek remains strong technical English/Chinese for coding but is not a full Work/Code shell. K3 vendor limitations call out sensitivity to thinking-history continuity and excessive proactiveness—constrain with system prompts / AGENTS.md if you need tight boundaries.

Speed. Flash is the throughput king in community reports. Pro is slower/more careful. K3 capacity is a different story—check current provider latency under load; early OpenRouter listings show single-provider Moonshot hosting at launch.

Community sentiment (Reddit / HN)

Kimi K3 praise: Open-weight frontier push at 2.8T; Hermes/OpenCode users finishing multi-step projects that stalled on weaker models; Arena/WebDev hype; AA top-tier intelligence and coding-agent index placements; Minecraft-scale long sessions as meme demos.

Kimi K3 complaints: “$3/$15 is not cheap Chinese AI anymore”—priced like frontier closed models; output-token burn on long max-thinking runs; context rot after ~300k in agent threads; local run impractical without massive RAM/VRAM; benchmaxxing skepticism until common harnesses land.

DeepSeek praise: V4 Flash as magical cheap coding implementor; Pro as diligent, cache-cheap reasoning; open MIT weights as exit hatch; HN obsession with sub-dollar multi-million-token burns and permanent Pro discounts.

DeepSeek complaints: Pro “underwhelmed” vs closed models or Kimi for some harnesses; Flash weak instruction-following on creative/roleplay presets; text-only gap; soft censorship / content policy debates; China-hosting + regulator scrutiny for sensitive orgs; peak-hour pricing uncertainty.

“V4 Flash for volume. V4 Pro or Kimi K3 when the task is actually hard. Route, don’t marry one logo.” — composite of OpenCode / Hermes / agent-blog consensus

When Kimi K3 wins

  • You need native vision or screenshot-in-the-loop coding/agents without bolting on a second model.
  • Long-horizon multi-step agents, Kimi Code CLI, Kimi Work research dashboards, or OpenClaw-style sustained sessions.
  • You want the current open-frontier ceiling (2.8T class, AA top-tier intelligence) and will pay Sonnet-class output rates for it.
  • Chinese-first or bilingual knowledge work and product UX matter more than pure $/MTok.
  • You prefer a product shell (membership, Work, Code) rather than only an API meter.
  • Cache-hit rates stay high on long coding prefixes so effective input cost drops toward $0.30/MTok.

When DeepSeek wins

  • Token bill dominates—batch jobs, evals, high-volume chat, cheap implementor loops on Flash.
  • You want plain MIT open weights available today and the cleanest redistribution/self-host story.
  • 1M context on the cheap SKU without jumping to $15/MTok output.
  • You already run a strong harness (OpenCode, Hermes, custom agent) and only need a model endpoint.
  • Speed/throughput (Flash) or decisive short coding turns beat multi-agent product chrome.
  • You need the cheapest serious text-only failover next to a closed planner (Claude/GPT).

Risks and failure modes

  • SKU confusion: Budgeting “DeepSeek” or “Kimi” without Flash/Pro or K3 ids blows TCO models.
  • Harness sensitivity: Coding scores swing with Kimi Code vs Claude Code vs OpenCode; K3 quality can degrade if thinking history is dropped mid-session.
  • K3 open-weights timing: API live first; full dump scheduled 2026-07-27—do not plan cluster deploys on vapor.
  • Jurisdiction & content policy: Both first-party APIs are China-based. Censorship/refusal patterns, privacy policies (PRC data residency), and regulator scrutiny apply—self-host or neutral hosts if policy requires.
  • Always-on agent risk (Kimi Claw / Work): Browser-native or file-system agents need explicit security review; forum reports include quota and destructive-action incidents.
  • K3 proactiveness: Vendor warns the model may improvise beyond tight boundaries—constrain system prompts.
  • Deprecations (DeepSeek): Migrate off legacy chat/reasoner model IDs before cutover dates.
  • Peak pricing / capacity: DeepSeek time-of-day surcharge discussions; K3 single-provider bottlenecks early after launch.

Recommendation by profile

ProfileDefault pickWhy
Indie shipping high volume CRUDDeepSeek V4 FlashLowest serious text cost; 1M ctx; good implementor loops
Hard multi-hour coding agentKimi K3 (or V4 Pro if budget-bound)Long-horizon + Code product; pay for quality sessions
Vision / UI screenshot loopsKimi K3Native multimodal; V4 is text-only first-party
Self-host / redistribute weights todayDeepSeek V4MIT weights live; K3 dump was still scheduled at research time
Regulated / data-residency strictSelf-host either or non-PRC hostFirst-party China hosting + scrutiny for both brands
CN-first product teamKimi K3 + productsWork/Code shell + bilingual strength
Hybrid production stackFlash implementor + K3/Pro plannerCommunity default pattern
Eval / benchmark farmDeepSeek V4 FlashCache + concurrency economics

FAQ

Is Kimi K3 better than DeepSeek V4 overall?
Not as a single number. K3 leads on open-frontier intelligence claims, multimodal, and productized long agents. V4 Flash leads on $/MTok and throughput; V4 Pro is the closer quality peer at a fraction of K3 output price. Task routing beats monogamy.

What does Kimi K3 cost on the API?
Official: $0.30 per 1M input tokens on cache hit, $3.00 on cache miss, $15.00 per 1M output, 1M context. Confirm on platform.kimi.ai pricing docs before budgeting.

What does DeepSeek V4 cost on the API?
Official: Flash $0.0028 / $0.14 / $0.28 (hit/miss/out); Pro $0.003625 / $0.435 / $0.87 per 1M. Both 1M context, 384K max output.

Are both open source?
DeepSeek V4 weights are MIT on Hugging Face now. Kimi K2-series is open under Modified MIT; K3 full weights were scheduled for July 27, 2026—API was live first. “Open” means weights + license, not necessarily free hosted inference.

Can I self-host Kimi K3?
After the public dump, in principle—expect multi-TB storage for usable quants and multi-accelerator serving. Many builders will keep using API/providers for months. DeepSeek V4 is already practical on third-party hosts and large local clusters.

Which is better for coding agents?
Split: Flash for cheap implementor loops; K3 (or Pro) for hard multi-step / repo-scale sessions. Harness choice (Kimi Code, OpenCode, Hermes) often moves the needle more than brand.

What about privacy and censorship?
Both first-party services are China-jurisdiction with documented content-policy and privacy concerns in press and research. Self-host open weights or use carefully vetted hosts for sensitive workloads.

Should I use both?
Yes for most production stacks: Flash for volume, K3/Pro for hard turns, closed planner if policy allows. Routing is the real architecture.

Sources

This comparison is backed by 150 unique source URLs in research_cache/kimi-k3-vs-deepseek_sources.json: official product/docs/pricing for both labs, GitHub/Hugging Face model cards, Artificial Analysis and independent reviews, Reddit and Hacker News threads (praise and complaints), security/censorship reporting, and multi-provider pricing pages. Citations like map to source ids in that file. Prices and model availability change—verify live vendor pages before production commits.

Bottom line

DeepSeek V4 is the default open-weight text workhorse of 2026: MIT weights today, absurd Flash economics, Pro when you need more head. Kimi K3 is the open-frontier bet: multimodal, long-horizon, productized agents, and Sonnet-class token pricing for sessions that justify it. Pick Flash for volume, K3 for hard multimodal agents, Pro when you want better DeepSeek quality without leaving the cheap family—and route between them instead of searching for a single winner.

Frequently Asked Questions

Is Kimi K3 better than DeepSeek V4?
No single winner. K3 leads on multimodal, long-horizon agents, and open-frontier intelligence. DeepSeek V4 Flash leads on cost and throughput; V4 Pro is the closer quality peer at lower output price.
How much does Kimi K3 API cost?
Official Moonshot pricing: $0.30 per 1M input tokens on cache hit, $3.00 on cache miss, $15.00 per 1M output, with a 1M-token context window.
How much does DeepSeek V4 API cost?
Official: V4 Flash $0.0028 cache-hit / $0.14 miss / $0.28 output per 1M; V4 Pro $0.003625 / $0.435 / $0.87. Both support 1M context and up to 384K max output.
Are Kimi K3 and DeepSeek V4 open source?
DeepSeek V4 Pro and Flash weights are on Hugging Face under MIT. Kimi K2-series uses Modified MIT; K3 full weights were scheduled for July 27, 2026, with API available first.
Which is better for coding agents?
Use DeepSeek V4 Flash for cheap implementor loops; use Kimi K3 or V4 Pro for hard multi-step repo work. Harness choice (Kimi Code, OpenCode, Hermes) often matters as much as the model.
Does Kimi K3 support vision?
Yes. K3 is natively multimodal. DeepSeek V4 first-party API is text-only, so screenshot-in-the-loop agents favor K3 unless you add a separate vision model.
Should I use both models?
Yes for most production stacks: Flash for volume text/coding, K3 or Pro for hard turns. Routing by task usually beats marrying one logo.

Intelligence Summary

The Final Recommendation

5/5 Confidence

Kimi K3 wins native vision, long-horizon agents, and open-frontier intelligence at $3/$15 miss/out ($0.30 cache-hit).

DeepSeek V4 Flash wins pure text cost ($0.14/$0.28 per 1M miss/out) with MIT open weights today.

Tool Profiles

Related Comparisons

Popular comparisons

Stay Informed

The Builder Switch Brief

When tools change pricing or features — plus the switch decisions that matter. Free.

Subscribe Free →