The study found that watermarking can change how AI models refuse harmful requests and select tools in agent systems. These behavioral effects, called sampling drift, vary by model and watermark key and can be hidden in aggregate performance scores.
log in to read full article