Scoring past conversations
An evaluation defines what you measure and how; running it produces a result. Scores accumulate against a stable definition, so trends mean something.
An evaluation scores conversations that already happened against criteria you
write down. It is how you find out what the agent is doing on real traffic,
none of which you scripted.
That split is deliberate. The definition is stable and version-controlled by
you; runs accumulate against it, so a score means something over time rather
than being a one-off measurement with its own private criteria.