Track use and improvement
Measure useful work without confusing requests with outcomes.
Track the actor, app, environment, workspace, operation/run ID and client surface. The same run may be opened through Brain, an API script and MCP; that is not automatically three scientific analyses.
Compare result quality, elapsed time, reliability, ease of use and cost across versions. Record the test inputs and scientific scope so comparisons are fair.
Label costs honestly: measured, estimated or unknown. Receiving a usage event does not by itself prove that all compute cost was captured. Reconcile worker completion, duplicate deliveries and retries before reporting totals.
Use the platform dashboard to inspect available observations. Keep run-level evidence so the team can understand why a number changed.
