Activation steering modifies inference-time behavior by adding or subtracting vectors in activation space.
Limitations
This often requires access to model weights and internals, which API-only usage may not allow.
Steer model behavior at inference time by manipulating internal activation patterns.
Activation steering modifies inference-time behavior by adding or subtracting vectors in activation space.
This often requires access to model weights and internals, which API-only usage may not allow.