Distinguishing between LLM injection detection methods
Distinguishing between LLM injection detection methods
In How might LLMs detect injected tokens? I described two methods that LLMs could use to detect injected tokens in their output:
How might LLMs detect injected tokens?
How might LLMs detect injected tokens?
Let's say Claude is generating some text in an autoregressive fashion. The output might look something like: