LLMgram · AI News · 2026-08-04

Redwood: subliminal fine-tune backdoor with ~100 teacher completions

Redwood: subliminal fine-tune backdoor with ~100 teacher completions

Redwood Research reports attackers can implant a covert backdoor by altering the teacher model for about 100 completions, roughly 0.5% of fine-tuning data, without needing prompt access. The result sharpens supply-chain risk for any lab that trusts small teacher-model deviations during fine-tuning.

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Attackers Can Subliminally Implant a Backdoor at Low Sample Count Without Prompt Access