Skip to main content
Rolling 7-day briefing

The LLM week, compressed.

A rolling 7-day briefing, distinct from today's Digest. Built from LLMgram's canonical AI Signal pipeline, ranked for source quality, event relevance, and usefulness to builders. Click any item to open its full AI Signal card without leaving LLMgram. For today's compressed MUST packet, open Today's Digest.

10signals selected
7dranking window
Sep 02, 2026 · 22:06 UTCgenerated
fresh sourceAI Signal data
Sep 02, 2026 · 22:06 UTCsource refreshed
Top 10 This Week
01
r/LocalLLaMA Top · model · Sep 01, 2026

New Model: Spark-X2.5-4B, Spark-X2.5-1.7B

I was browsing HF for small LLMs and run into this model. It does not seem to be a fine tune - the model has its own architecture. https://huggingface.co/XHToken/Spark-X2.5-1.7B https://huggingface.co/XHToken/Spark-X2.5-4B There are 4B/1.7…

Builder angle:

02
Techmeme · model · Sep 02, 2026

Muse Spark 1.3 with max reasoning, in limited preview for partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Fable 5.1 and Opus 5 (…

@artificialanlys : Muse Spark 1.3 with max reasoning, in limited preview for partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Fable 5.1 and Opus 5 — Meta has released Muse Spark 1.3, their fourth Muse Spark mo…

Builder angle:

03
MarkTechPost · model · Sep 02, 2026

Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber: One Core Model, Two Access Envelopes

Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2, 2026. Both variants run on the same foundational intelligence, split by safety mitigations rather than model size. Gemini 3.8 Flash is generally available at $0.75…

Builder angle:

04
The Decoder · model · Sep 02, 2026

Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA

Google's Gemini 3.8 Flash, the third Flash model in six weeks, matches Claude Opus 5 on some agentic coding benchmarks at lower cost. But its "working harder" reasoning burns about 30 percent more output tokens per task, making it pricier…

Builder angle:

05
Towards AI · model · Sep 02, 2026

Why I’d Swap Claude for Qwen3.8-Max

Claude Fable 5 wins the benchmarks. Qwen3.8-Max costs $0.91 per task to Fable’s $3.14. That’s 3 attempts against 1… let your test suite… Continue reading on Towards AI »

Builder angle:

06
Vercel AI · model · Sep 02, 2026

GLM-5.3 is 50% off through DigitalOcean on AI Gateway

GLM-5.3 is 50% off on AI Gateway through Tuesday, September 8, in partnership with DigitalOcean. How to use the model during the offer period Using the promo name ( zai/glm-5.3-promo-50 ) gets the discounted rate. It routes only to Digital…

Builder angle:

07
smol.ai AI News · model · Sep 01, 2026

Claude Fable 5.1 and Claude Mythos 5.1

**Anthropic** released **Claude Fable 5.1** and **Claude Mythos 5.1**, which share base weights but differ in safeguards and routing, showing improved coding performance and usability with a **75% cache-read price cut to $0.25/MTok**. Benc…

Builder angle:

08
arXiv cs.AI · model · Sep 02, 2026

Asymmetries in Spontaneous and Instructed Deception

arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship be…

Builder angle:

09
arXiv cs.LG · model · Sep 02, 2026

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

arXiv:2609.00363v1 Announce Type: new Abstract: Conformance suites for quantized GEMM kernels ask whether two implementations agree within a tolerance. We measure what such a suite can detect. Injecting nine faults into a reference INT8 pi…

Builder angle:

10
arXiv cs.CL · model · Sep 01, 2026

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5…

Builder angle: