evalr.measures¶
End-to-end measures, defined generically: task completion, drop-off and rewrites.
Task completion is judged: a verdict type for any evaluator (a DSPy judge or a decision model)
reading a Transcript, or given by people. Drop-off and rewrites are computed from recorded
activity, exactly and cheaply. The libraries' [evals] extras put their logs into these
inputs.
See Workflow measures.
Inputs¶
What the measures read, in terms any application's log can be put in.
Session
pydantic-model
¶
Activity
pydantic-model
¶
Bases: BaseModel
Something that happened in a session.
Attributes:
| Name | Type | Description |
|---|---|---|
at |
datetime
|
When. |
role |
Role
|
Who did it. |
kind |
str
|
What: |
ref |
str | None
|
Pairs a proposal with its resolution. |
Fields:
History
pydantic-model
¶
Revision
pydantic-model
¶
Transcript
pydantic-model
¶
Role
¶
Role = Literal['person', 'agent', 'system']
Who acted: a person, an agent, or the system itself.
Task completion¶
Judged by any evaluator, or given by people.
TaskCompletion
pydantic-model
¶
completion_rate
¶
completion_rate(
verdicts: Iterable[
Verdict[TaskCompletion] | TaskCompletion
],
) -> float | None
The share of sessions completed; None for none.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdicts
|
Iterable[Verdict[TaskCompletion] | TaskCompletion]
|
One per session: an evaluator's verdict, or a plain value, such as a person's feedback. |
required |
Drop-off¶
Computed from a session's activity.
DropOff
pydantic-model
¶
Bases: BaseModel
Whether a session ended with the agent waiting on a person who never came back.
Fields:
measure_drop_off
¶
Decide whether a session dropped off.
A session dropped off when the agent acted last and no person acted within window, or
when a proposal was left unresolved for longer than window. Until the window has passed,
it is pending. A session whose last word was a person's continued.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
session
|
Session
|
The session's activity. |
required |
window
|
timedelta
|
How long a person has to come back. |
required |
now
|
datetime
|
When the log was read: the end of the recording. |
required |
drop_off_evaluator
¶
drop_off_evaluator(
*, window: timedelta, now: datetime
) -> FunctionEvaluator[Session, DropOff]
Drop-off as an evaluator, versioned by its window.
drop_off_rate
¶
Rewrites¶
Computed from an artifact's revisions.
Rewrites
pydantic-model
¶
measure_rewrites
¶
Count the agent's revisions that a person substantially rewrote within a window.
A person's revision rewrites the agent's when it comes within window after it, before
the agent writes again, and changes at least threshold of its text, as
share_changed measures it. Of several such revisions, the last counts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
history
|
History
|
The artifact's revisions. |
required |
window
|
timedelta
|
How soon after the agent a person's change counts. |
required |
threshold
|
float
|
The share of the text a change must alter to count as a rewrite. |
0.2
|
share_changed
¶
The share of a text that an edit changed, from 0 (none of it) to 1 (all of it).
It is one minus the texts' similarity: twice the characters they share, over the characters
in both. It aligns lines, then compares characters within the lines that changed, with
difflib's autojunk heuristic off. On a text of 200 characters or more, that heuristic
ignores every character that makes up more than 1% of it, which calls a small change to a
table or a checklist a rewrite.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
before
|
str
|
The text as it was. |
required |
after
|
str
|
The text as it became. |
required |
rewrite_evaluator
¶
rewrite_evaluator(
*, window: timedelta, threshold: float = 0.2
) -> FunctionEvaluator[History, Rewrites]
Rewrites as an evaluator, versioned by its window and threshold.
The leading number is the measure's own version, bumped when its computation changes.