
Watermarking an AI model is supposed to make its outputs easier to identify without changing what the model does. But new research from Lasso Security suggests that the picture becomes more complicated when those same models are used inside AI agents.
Lasso tested Google DeepMind’s SynthID-Text watermarking method across several open models and found that enabling the watermark could change which tools models selected, what arguments they passed to those tools, and whether they refused harmful requests. In some cases, these changes became much larger under prompt injection.
Why should you care? Anthropic announced in August that future Claude models will use a version of SynthID-Text. The company says the watermark has no practical effect on the quality or content of its outputs. The EU AI Act also requires providers of systems that generate synthetic text and other media to make their outputs machine-readable and detectable as AI-generated or manipulated.
Lasso’s findings do not dispute SynthID’s ability to preserve text quality. But the researchers look at a different property: whether adding the watermark changes individual model decisions. They call this effect “sampling drift.”
“Our findings support evaluating provenance and behavioral reliability together,” Andrea Siposova, AI Security Researcher at Lasso, told TechTalks.
How watermarks cause sampling drift
A text watermark creates a statistical pattern in the model’s output that can later be detected. SynthID-Text inserts the signal while the model generates tokens rather than adding it to the completed response as metadata or hidden characters.
To understand why that can change behavior, it helps to look at how language models generate text.
At each step, a language model calculates probabilities for the tokens that could come next. The model might assign very high probability to one token, or it might consider several alternatives plausible. A sampling algorithm then uses those probabilities to select the next token.
SynthID-Text modifies this process through a technique called “Tournament sampling.” It samples several candidate tokens from the model’s original probability distribution and puts them through rounds of competition. Pseudorandom scores derived from a watermark key and the recent text help determine which candidates advance. The winner becomes the next generated token. Repeating this process creates a statistical pattern that can later be detected with the appropriate key.
SynthID can be configured to preserve the model’s underlying token distribution when choosing the watermark. Google DeepMind reported no measurable reduction in text quality in an evaluation covering nearly 20 million responses from live Gemini interactions.
But preserving the overall distribution does not mean that a specific prompt will produce exactly the same sequence of tokens with and without a watermark. A particular watermark key can cause a different candidate to win the tournament.
This difference becomes more important when the model is uncertain about what comes next. In machine learning, this uncertainty is often described as “entropy.” Low entropy means one or a few tokens dominate the probability distribution. High entropy means there are several plausible choices.
“The original SynthID-Text paper explains that tournament sampling embeds a stronger watermark when there is more entropy in the model’s next-token distribution,” Siposova said. “Consistent with this, our token-level analysis found that watermarking more frequently changed the highest-probability token at higher-entropy positions.”
For normal prose, changing one plausible token for another might only alter the wording of a sentence. Inside an AI agent, the same token can be part of an amount, query, file path, date, recipient, or resource identifier. A small change in generation can therefore become a different action.
There is no simple way to identify these vulnerable positions beforehand. “However, this does not provide a specific way to predict vulnerable fields from a JSON schema alone,” Siposova said.
The same model can make different decisions
Lasso tested this effect through paired experiments. Each task was run with and without SynthID while keeping the model, seed, batch composition, order, and temperature the same.
For tool use, the researchers used BFCL v4, a benchmark that tests whether models can select the correct tool and generate the appropriate arguments. For safety behavior, they tested 200 harmful requests from HarmBench and 100 benign controls from JailbreakBench. The harmful requests were tested normally and with a fixed prompt-injection technique.
On tasks where the model was expected to make a tool call, watermarking reduced accuracy on six of the seven tested models. Four of the decreases were statistically significant.
But the aggregate accuracy figures hide a more interesting pattern. Suppose two versions of a model answer 100 tasks and both get 90 right. They can have identical accuracy even if the watermarked version gets one previously correct task wrong and fixes another task that was previously wrong.
Lasso measures these changes using “churn,” or the paired disagreement rate: the percentage of tasks whose correctness verdict changes between the watermarked and unwatermarked runs.
Across 21 combinations of models and temperature settings, average churn was 6.5%. In some cases, the gap between churn and the headline accuracy change was much larger. For example, at temperature 1.0, 16.8% of phi-4’s individual tool-call verdicts changed, while its overall accuracy fell by only 2.87 percentage points. Llama-3.1-8B had 9.9% churn while losing just 0.87 accuracy points.
From a practical standpoint, the two measurements answer different questions. Aggregate accuracy shows how often the model succeeds overall. Churn shows how frequently changing the deployment configuration alters the outcome on individual requests.
“It would be helpful for developers if providers published capability and safety evaluations when introducing or updating watermarking, similar to the evaluations published with new models,” Siposova said. “These could cover tool calling, instruction following, and safety behavior, with comparisons against the previous configuration.”
Like benchmarks used when choosing a model, these results would help developers decide what to assess before deployment.
“They, however, would not predict the full effect on a specific application, because agent behavior also depends on its harness and other behavior-defining elements (prompts, tool definitions, orchestration, and execution controls),” she added.
Different models fail differently
Lasso divided tool-call failures into incorrect tool selection, incorrect arguments, and malformed output. Llama-3.1-8B’s accuracy loss came mostly from incorrect arguments, which accounted for a 3.48-point decline, followed by wrong-tool calls at 1.84 points.
For phi-4 and Granite-3.2-8B, malformed output was the main source of deterioration, accounting for declines of 5.96 and 4.36 points respectively.
A malformed tool call will often fail parsing and never execute. A valid call containing the wrong file path, amount, recipient, or resource identifier can pass basic structural checks and perform an unintended action.
“Malformed calls generally fail before execution, so the priority is reliable error handling and validated retries,” Siposova said. “Schema validation can catch structural errors in parseable calls, but it cannot establish whether argument values match the intended task.”
Prompt injection amplified some changes
Lasso also tested whether sampling drift affects refusal behavior.
The researchers first gave the models harmful requests and measured whether they complied or refused. They then repeated the experiment after adding one fixed prompt-injection technique presented as retrieved content and designed to weaken the model’s refusal.
At a temperature of 0.001, Gemma-3-27B had 6% churn between watermarked and unwatermarked responses on the original harmful prompts. Under prompt injection, churn rose to 23.5%. The net change in harmful compliance shifted from a one-point decrease to a 12.5-point increase.
If you enjoyed this article, please consider supporting TechTalks with a paid subscription (and gain access to subscriber-only posts)
Gemma-3-12B showed a similar pattern. Churn increased from 7.5% to 11%, while the net effect changed from a 0.5-point reduction in harmful compliance to a nine-point increase.
Other models behaved differently. Phi-4 and Qwen3-4B changed little under either condition, although both already tended to refuse benign requests in the evaluation. Lasso therefore warns against interpreting their stability as evidence that watermarking preserves safety behavior better on those models.
The results also do not establish that watermarking generally pushes models toward harmful compliance. “This describes the pattern in our tested sample, rather than a general rule about watermarking,” Siposova said.
The watermark key changes the outcome
SynthID uses the key and recent token context to generate the pseudorandom scores that influence Tournament sampling. Lasso repeated its prompt-injection experiment using 11 different keys and found that changing the key could change both the size and direction of the effect.
For Llama-3.1-8B, for example, the ten additional keys produced changes in attack success ranging from a 4.5-point decrease to a 14.5-point increase relative to the unwatermarked model. Granite-3.2-8B also moved in both directions depending on the key.
Siposova says there is a plausible explanation, though the experiments do not prove it.
“One hypothesis is that, when both refusal and compliance continuations are plausible, watermarking could shift token selection toward either one,” she said. “The selected token then becomes part of the context for subsequent generation and could steer the rest of the response.”
Once a token is generated, the model sees it as part of the context when choosing the next token. An early change can therefore send the rest of the response down a different path.
“Our finding of greater token-level sensitivity at higher entropy is consistent with this possibility, but does not demonstrate that mechanism,” Siposova said.
Treat watermarking as a deployment change
The findings suggest that developers should treat watermarking as another behavior-defining part of an AI application rather than assuming it is invisible to the system’s operation.
“The practical first step is to run behavioral assessment of the application under the new deployment configuration, similarly as they would when changing models,” Siposova said.
That means rerunning application-specific tests against the configuration that will actually be deployed. For tool-using agents, developers should separately measure malformed calls, incorrect tool selection, and incorrect arguments instead of relying on one aggregate accuracy score.
Siposova recommends paying particular attention to values whose exact contents determine what an application does, such as amounts, quantities, dates, file paths, and record or resource identifiers.
Runtime controls also need to account for failures that cannot be caught through schema validation.
For parseable tool calls, Siposova recommends using pre-execution hooks or middleware to compare arguments against trusted records, explicit user input, and application state. When an application already knows a value, it should supply that value directly rather than asking the model to generate it. Identities, credentials, and permissions should come from the authenticated session and be enforced by the agent harness.
For consequential actions that cannot be independently checked, a preview or user confirmation can provide another layer of protection.
“The aim is to separate the model’s proposed action from the application’s decision to execute it,” Siposova said.
Developers working with black-box model APIs face an additional problem. They might be able to observe a behavioral change without knowing whether the provider changed its watermarking configuration or rotated its key.
Regular regression tests and safety evaluations can help. Runtime monitoring of tool calls, validation failures, retries, execution results, and user corrections can also reveal shifts in production behavior. But these signals cannot reliably identify a watermark change as the cause.
Providers can make that process easier by publishing capability and safety evaluations when they introduce or modify watermarking and by documenting changes to the provenance method and configuration.























