Skip to content

evalr.measures

End-to-end measures, defined generically: task completion, drop-off and rewrites.

Task completion is judged: a verdict type for any evaluator (a DSPy judge or a decision model) reading a Transcript, or given by people. Drop-off and rewrites are computed from recorded activity, exactly and cheaply. The libraries' [evals] extras put their logs into these inputs.

See Workflow measures.

Inputs

What the measures read, in terms any application's log can be put in.

Session pydantic-model

Bases: BaseModel

A conversation or a causal chain, as a timeline of activity.

Fields:

activities pydantic-field

activities: list[Activity]

Everything that happened, oldest first

Activity pydantic-model

Bases: BaseModel

Something that happened in a session.

Attributes:

Name Type Description
at datetime

When.

role Role

Who did it.

kind str

What: message by default, proposal for a change awaiting a person's decision, and resolution for that decision. Other kinds count as activity.

ref str | None

Pairs a proposal with its resolution.

Fields:

History pydantic-model

Bases: BaseModel

The revisions of one artifact, oldest first.

Fields:

Revision pydantic-model

Bases: BaseModel

One version of something an agent and people both write, such as an artifact.

Attributes:

Name Type Description
at datetime

When it was written.

role Role

Who wrote it.

text str

Its whole content, as text.

Fields:

Transcript pydantic-model

Bases: BaseModel

What a judge reads to decide whether a task was completed.

Fields:

request pydantic-field

request: str

What the person asked for

turns pydantic-field

turns: list[Turn]

The conversation, oldest first

result pydantic-field

result: str | None = None

What was produced: the final artifact or report, as text

Turn pydantic-model

Bases: BaseModel

One turn of a conversation.

Fields:

Role

Role = Literal['person', 'agent', 'system']

Who acted: a person, an agent, or the system itself.

Task completion

Judged by any evaluator, or given by people.

TaskCompletion pydantic-model

Bases: BaseModel

Whether the session achieved what the person asked for, and how well.

Fields:

completed pydantic-field

completed: bool

The person's request was achieved

quality pydantic-field

quality: int

How well it was done, from 1 (poorly) to 5 (very well)

reason pydantic-field

reason: str | None = None

Why, in a sentence

completion_rate

completion_rate(
    verdicts: Iterable[
        Verdict[TaskCompletion] | TaskCompletion
    ],
) -> float | None

The share of sessions completed; None for none.

Parameters:

Name Type Description Default
verdicts Iterable[Verdict[TaskCompletion] | TaskCompletion]

One per session: an evaluator's verdict, or a plain value, such as a person's feedback.

required

Drop-off

Computed from a session's activity.

DropOff pydantic-model

Bases: BaseModel

Whether a session ended with the agent waiting on a person who never came back.

Fields:

  • outcome (Literal['continued', 'dropped', 'pending'])
  • cause (Literal['no_reply', 'unresolved_proposal'] | None)

outcome pydantic-field

outcome: Literal['continued', 'dropped', 'pending']

continued: a person acted after the agent's last turn, within the window; dropped: nobody did, or a proposal was left unresolved; pending: too soon to tell

measure_drop_off

measure_drop_off(
    session: Session, *, window: timedelta, now: datetime
) -> DropOff

Decide whether a session dropped off.

A session dropped off when the agent acted last and no person acted within window, or when a proposal was left unresolved for longer than window. Until the window has passed, it is pending. A session whose last word was a person's continued.

Parameters:

Name Type Description Default
session Session

The session's activity.

required
window timedelta

How long a person has to come back.

required
now datetime

When the log was read: the end of the recording.

required

drop_off_evaluator

drop_off_evaluator(
    *, window: timedelta, now: datetime
) -> FunctionEvaluator[Session, DropOff]

Drop-off as an evaluator, versioned by its window.

drop_off_rate

drop_off_rate(
    verdicts: Iterable[Verdict[DropOff] | DropOff],
) -> float | None

The share of decided sessions that dropped off; None if none is decided.

Parameters:

Name Type Description Default
verdicts Iterable[Verdict[DropOff] | DropOff]

One per session: an evaluator's verdict, or a plain value, such as one from measure_drop_off.

required

Rewrites

Computed from an artifact's revisions.

Rewrites pydantic-model

Bases: BaseModel

How much of what an agent wrote people substantially rewrote.

Fields:

agent_revisions pydantic-field

agent_revisions: int

Revisions the agent wrote

rewritten pydantic-field

rewritten: int

Of those, the ones people rewrote

rate pydantic-field

rate: Annotated[float, Field(ge=0.0, le=1.0)] | None = (
    None
)

Rewritten over agent revisions; empty without any

measure_rewrites

measure_rewrites(
    history: History,
    *,
    window: timedelta,
    threshold: float = 0.2,
) -> Rewrites

Count the agent's revisions that a person substantially rewrote within a window.

A person's revision rewrites the agent's when it comes within window after it, before the agent writes again, and changes at least threshold of its text, as share_changed measures it. Of several such revisions, the last counts.

Parameters:

Name Type Description Default
history History

The artifact's revisions.

required
window timedelta

How soon after the agent a person's change counts.

required
threshold float

The share of the text a change must alter to count as a rewrite.

0.2

share_changed

share_changed(before: str, after: str) -> float

The share of a text that an edit changed, from 0 (none of it) to 1 (all of it).

It is one minus the texts' similarity: twice the characters they share, over the characters in both. It aligns lines, then compares characters within the lines that changed, with difflib's autojunk heuristic off. On a text of 200 characters or more, that heuristic ignores every character that makes up more than 1% of it, which calls a small change to a table or a checklist a rewrite.

Parameters:

Name Type Description Default
before str

The text as it was.

required
after str

The text as it became.

required

rewrite_evaluator

rewrite_evaluator(
    *, window: timedelta, threshold: float = 0.2
) -> FunctionEvaluator[History, Rewrites]

Rewrites as an evaluator, versioned by its window and threshold.

The leading number is the measure's own version, bumped when its computation changes.

rewrite_rate

rewrite_rate(
    verdicts: Iterable[Verdict[Rewrites] | Rewrites],
) -> float | None

Rewritten over agent revisions across artifacts; None without any agent revision.

Parameters:

Name Type Description Default
verdicts Iterable[Verdict[Rewrites] | Rewrites]

One per artifact: an evaluator's verdict, or a plain value, such as one from measure_rewrites.

required