DeepSeek Resumes $8 Billion Round With Monolith in the Running
Bloomberg, citing people familiar, says DeepSeek resumed a second round seeking about eight billion dollars with Monolith Management in talks, after talks paused when Liang Wenfeng's investor remarks leaked. Techmeme relayed Bloomberg reporting of a roughly seventy-four billion dollar valuation. The Information reports DeepSeek is planning model price increases while restarting funding talks. A LocalLLaMA user says DeepSeek-V4-Flash-0731 is fully broken on one AMD MI325X with vLLM 0.26.0, while another post shows fifty-four gigabyte IQ2_XXS quantization reaching about twenty point five tokens per second. Bindu Reddy on X cites GPU shortages forcing hosts to turn off DeepSeek Flash for latency. Production teams face mixed signals between capital backing and real inference friction, and Bloomberg leaves Monolith's stake and closing timing unconfirmed.
DeepSeek Resumes $8 Billion Round With Monolith in the Running
DeepSeek has resumed its second funding round, seeking close to $8 billion, with Monolith Management in talks to contribute, according to people familiar with the matter.
Key takeaway
DeepSeek is pairing a near-$8 billion capital raise with planned price increases, betting scaled funding can sustain its model stack despite uneven local deployment and host-side GPU bottlenecks.
What happened
According to Bloomberg, citing people familiar with the matter, DeepSeek has resumed its second funding round seeking close to $8 billion, with Monolith Management in talks to contribute. Techmeme attributed reporting to Bloomberg says the company is targeting the round at roughly a $74 billion valuation after pausing talks following a leak of founder Liang Wenfeng's investor remarks.
The Information reports DeepSeek is simultaneously planning to hike model prices as it resumes funding talks. On the deployment side, a r/LocalLLaMA user says DeepSeek-V4-Flash-0731 runs completely broken on a single AMD Instinct MI325X with vLLM 0.26.0, while another post documents a 54GB IQ2_XXS quantized build reaching about 20.5 tokens per second.
Evidence
DeepSeek resumed its second funding round seeking close to $8 billion with Monolith Management in talks to contribute
Bloomberg Technology · attributed
DeepSeek has resumed its second funding round, seeking close to $8 billion, with Monolith Management in talks to contribute, according to people familiar with the matter.
DeepSeek is seeking $8B at a $74B valuation after pausing talks following a leak of Liang Wenfeng's remarks to investors
Techmeme · attributed
Sources: DeepSeek has resumed its funding round, seeking $8B at a $74B valuation, after pausing talks following the leak of Liang Wenfeng's remarks to investors (Bloomberg)
DeepSeek is resuming funding talks while planning to increase model pricing
The Information AI · attributed
According to The Information, DeepSeek is resuming funding talks while simultaneously planning to increase its model pricing.
DeepSeek-V4-Flash-0731 runs completely broken on AMD MI325X with vLLM 0.26.0
r/LocalLLaMA Top · attributed
A user reports that DeepSeek-V4-Flash-0731 runs 'completely broken' on a single AMD Instinct MI325X using vLLM 0.26.0, despite correct tokenizer and parser configurations.
A community DeepSeek V4 Flash build using IQ2_XXS quantization reaches about 20.5 tokens per second at 54GB
r/LocalLLaMA Top · attributed
A community-driven adaptation of DeepSeek-V4-Flash utilizes extreme quantization (IQ2_XXS) to reduce the model size from 95GB to 54GB, achieving ~20.5 tokens per second.
GPU supply constraints are forcing hosts to turn off DeepSeek Flash due to latency
Bindu Reddy (X) · attributed
We are officially running out of GPUs to host these models and demand is out-stripping supply by a mile Having to turn off DeepSeek Flash because it's so slow right now
Why it matters
Builders weighing DeepSeek for production must reconcile bullish capital-market backing with real-world inference friction on AMD stacks, quantization tradeoffs, and third-party hosting capacity.
Limits and uncertainties
Bloomberg cites people familiar with the matter; Monolith participation and final round size are not publicly confirmed.
The MI325X vLLM failure is reported by a single LocalLLaMA user on one hardware and software configuration.
The Information pricing and funding details are not independently confirmed in the Bloomberg report included here.
Practical implications
Pin and test vLLM and ROCm builds before committing AMD MI325X clusters to DeepSeek V4 Flash workloads.
Budget for potential DeepSeek API price increases when modeling inference cost projections.
Plan fallback hosting paths if third-party open-weight inference faces GPU capacity constraints.
What to watch
Whether DeepSeek closes the reported $8 billion round and at what valuation.
vLLM release notes and community fixes for DeepSeek-V4-Flash-0731 on ROCm and MI325X.
Official DeepSeek pricing changes tied to resumed funding talks.