Skip to content

Local AI

I Put DeepSeek V4 Flash on Two Lunchbox-Sized Computers. Then I Tried to Break It.

A preregistered local deployment study of DeepSeek-V4-Flash-0731 on a dual-DGX-Spark cluster found a practical tie with Qwen, severe effective-context limits, and blocked frontier comparisons.

TL;DR

  • Bottom line: DS4 on two Sparks was useful, but it did not beat the prior Qwen local result in a defensible way.
  • DS4 had zero primary practical execution errors, but more truncations. The practical result is a tie within the approved evidence, not a frontier-comparison completion.
  • Advertised context was not effective context. The study produced one scored 32K exact-fit slot and fourteen tokenization or capacity failures before health stop.
  • Human review is still 0/2. Non-human sensitivity is secondary, and hosted/frontier/public-agentic comparisons remain blocked.

The Short Version

I put DeepSeek-V4-Flash-0731 on a local dual-DGX-Spark cluster. The official DeepSeek model card describes it as a 284B-parameter model (pinned model card), which is the only place this numeric hook appears in the article. The root checkpoint is the official DeepSeek repository for DeepSeek-V4-Flash-0731 (pinned model card, model README). The hardware class is NVIDIA DGX Spark, a compact workstation line documented by NVIDIA (specifications, guide).

The result is more interesting than a win/loss headline.

DS4 passed 34 of 80 local practical slots. Qwen passed 35 of 80 paired slots. DS4 had zero execution errors in the primary practical analysis, which matters. It also had 14 truncations, twice Qwen's 7. So the practical read is a tie with different failure shapes.

No SOTA claim. No hosted frontier claim. No Terminal-Bench claim. No RULER claim. Terminal-Bench 2.1 and RULER are real external benchmark surfaces (Terminal-Bench repository, Terminal-Bench release note, RULER repository, RULER paper), but this study did not complete those tracks.

What The Evidence Actually Says

Evidence coverage

What is supported, qualified, or blocked.

supported

Practical DS4

80 planned local DS4 slots

claim-practical-ds4-primary

supported

Practical Qwen

80 paired Qwen slots

claim-practical-qwen-primary

qualified

Systems

30 supplemental concurrency records

claim-systems-supplemental

supported

Long context

1 scored exact-fit, 14 failures

claim-long-context-capacity-negative

blocked

Hosted/frontier

Hosted fidelity, frontier controls, Terminal-Bench blocked

claim-no-terminal-bench-frontier-power

blocked

Human review

0 of 2 independent human reviews complete

claim-llm-sensitivity-secondary

Source: claim ledger and publication-data hash b196a08c2cb5.

The controlled evidence is useful, but bounded. This is a preregistered local deployment study with material execution limitations, not a frontier-comparison completion.

Practical Tie, Different Failure Shape

Practical comparison

A practical tie, not a frontier win.

DS4 local dual-Spark

pass
34
fail
32
error
0
truncation
14

Qwen3.8 local prior study

pass
35
fail
29
error
9
truncation
7
Paired DS4 versus Qwen discordance counts
Pair classSlots
both nonpass34
both pass23
ds4 nonpass qwen pass12
ds4 pass qwen nonpass11

Delta DS4 minus Qwen: -0.0125 with 95% interval -0.1375 to 0.1125. Errors and truncations remain in denominators.

My read: DS4 was operationally cleaner in one important way, zero execution errors. But the extra truncations are not cosmetic. A truncation is a failed delivery state when the task needs a complete artifact.

The Systems Path Bent Under Load

Systems envelope

Throughput fell as concurrency rose.

Concurrency levels

28.5

aggregate output tok/s

27.4

median per-record tok/s

404.7

median TTFT ms

3,101.1

median wall ms

Supplemental systems metrics only. No power, energy, or soak reliability claim.

The serving path did useful work. It also showed a clear systems envelope. These numbers are supplemental only. They do not support energy, power, or cost claims.

The Context Cliff

Context cliff

Advertised capacity was not effective scored context.

Planned slots
15
Scored exact-fit
1
Capacity failures
14

The one scored slot had 32,765 filled prompt tokens. Health stop: activated; final soak blocked_not_run_due_health_degradation.

The endpoint advertised a large model length in the public environment record, but the study only scored one exact-fit 32K slot. The remaining long-context slots ended as tokenization or capacity failures. The final health stop prevented turning a degraded service into a fake reliability story.

Failures Worth Keeping

Failure taxonomy

The failures are part of the result.

model-output failure61
truncation/output-budget21
tokenization/capacity failure14
harness defect corrected by amendment10
provider/normalization error9
endpoint identity/credential/quota block7
environment prerequisite block1
health-stop/soak block1

Source: failure-taxonomy.json. Private joins and blind identity mapping are excluded.

Claim Status

Claim status

No statement exceeds the ledger.

supported: 6qualified: 4blocked: 4
supportedclaim-practical-ds4-primary

Local DS4 primary practical analysis covers 80 planned slots with amended validator revision applied.

supportedclaim-practical-qwen-primary

Qwen v2 primary practical analysis covers 80 planned slots with validator revision sidecar present for all rows.

supportedclaim-no-composite

No post-hoc weights or composite intelligence score are computed.

qualifiedclaim-llm-sensitivity-secondary

Completed non-human LLM sensitivity ratings and non-human adjudication are included only as secondary sensitivity analysis with 5 adjudicated sample-dimensions; public pooled scores are shown with DS4/Qwen subject breakdowns and are not human-review authority.

supportedclaim-paired-discordance

Paired DS4-vs-Qwen outcomes are computed only on the 80 identical task/repetition slots; discordance counts are {'both_pass': 23, 'both_nonpass': 34, 'ds4_pass_qwen_nonpass': 11, 'ds4_nonpass_qwen_pass': 12}.

qualifiedclaim-deployment-model-identity

Public environment/preflight evidence observed requested and returned model deepseek-v4-flash-0731 with verified_live status; this is a deployment observation, not a scored quality claim.

qualifiedclaim-deployment-topology-capacity

Public environment/preflight evidence observed a head API role, a second TP=2 worker role, advertised max_model_len=1048576, vLLM 0.25.2.dev0+g752a3a504.d20260714, cache dtype nvfp4_ds_mla, and KV cache size 1203768; these are observations, not effective-context or vendor benchmark claims.

qualifiedclaim-systems-supplemental

Systems metrics summarize 30 actual v2 concurrent records by concurrency; metrics are supplemental and do not support power or energy claims.

supportedclaim-long-context-capacity-negative

Long-context planned denominator is 15 slots: 1 exact-fit scored slot and 14 observed tokenization/capacity failures.

blockedclaim-no-long-context-quality-above-32k

No claim is made for retrieval/synthesis quality above the one 32K exact-fit scored slot.

blockedclaim-no-soak-reliability

No two-hour soak reliability claim is available because the health-stop rule blocked final soak.

blockedclaim-no-hosted-fidelity

No hosted fidelity claim is available because hosted DS4/control tracks were blocked.

blockedclaim-no-terminal-bench-frontier-power

No Terminal-Bench, frontier-control, hosted-fidelity, power, energy, or 2h-soak claims are included.

supportedclaim-strong-negative-failures

Failure taxonomy records strong negative findings across categories: {'endpoint identity/credential/quota block': 7, 'environment prerequisite block': 1, 'harness defect corrected by amendment': 10, 'health-stop/soak block': 1, 'model-output failure': 61, 'provider/normalization error': 9, 'tokenization/capacity failure': 14, 'truncation/output-budget': 21}.

Every claim above is bounded by the ledger. Human review remains 0/2. Non-human sensitivity is secondary. Hosted DS4 fidelity, public-agentic comparisons, frontier controls, Terminal-Bench, RULER, power, energy, and two-hour soak reliability are not available.

Sources