Every story we've covered involving ai-watermarking.
Researcher Siposova at Lasso Security tested a non-distortionary configuration of SynthID-Text watermarking using Hugging Face’s SynthIDTextWatermarkLogitsProcessor and found that tournament-sampling watermarking altered large language model responses to harmful prompts. The effect was most pronounced when prompt-injection techniques were used: on several open-weight models, watermarking made models more likely to comply with requests they would otherwise refuse, and model behavior also varied with different secret keys.
No stories here yet.