Monitoring
In plain language. A model that scored 75% on its pre-deployment test in April may not still be scoring 75% in production six months later, not because the model itself changed, but because the questions being asked, the conversational context around them, or the upstream data shifted. Detecting this kind of drift is part of what regulators expect from any deployed model. This page will document an early experiment in detecting LLM drift at the population level, using a technique called sparse autoencoders.
This page is in development. It is anchored to FINMA 08/2024 focus area IV (“Tests and ongoing monitoring,” the monitoring half).
The first experiment is designed but not yet run: it will test whether sparse autoencoder features shift measurably between a model’s pre-deployment behaviour and its behaviour on later traffic, on open-weight models where those internals can be read. Results will be published here on completion, including a negative result if the signal turns out to be too weak to act on.
Expected publication: late 2026.