V7 Go added organisation wide reporting so metrics can be read across every workspace rather than one at a time, began recording manual field edits with a timestamp and surfacing them in token reports, and started flagging output samples as stale when they no longer match the instructions that generated them.
Our readTracking manual edits is the quiet one and it is graded on oversight for a reason. A record of every point where a human corrected the model is the raw material for measuring where the system is actually unreliable, and most platforms in this index cannot produce it.