Skip to content

evalr.langfuse

Langfuse datasets, scores, score configs and experiments (the [langfuse] extra).

Langfuse holds working datasets and experiments, next to the traces they judge (ADR-0003), and the scores of evaluators and people alike (ADR-0012). This package adapts it to the DatasetStore, ScoreSink, ScoreConfigStore and ExperimentTracker ports (ADR-0006), through the application's own Langfuse client.

See Langfuse datasets, Scores in Langfuse and Experiments in Langfuse.

Datasets

LangfuseDatasetStore

LangfuseDatasetStore(client: Langfuse)

Saves datasets to Langfuse, and loads them from it, through a Langfuse client.

The client is the application's, configured as it likes; the store makes no other connection. Its calls are synchronous, so they run in a worker thread.

Use a Langfuse client.

Parameters:

Name Type Description Default
client Langfuse

The application's client.

required

save async

save(dataset: Dataset[InputT, VerdictT]) -> str

Create or update the Langfuse dataset, item by item, and archive what was removed.

Only items whose content changed are written.

Returns:

Name Type Description
str

A time at which Langfuse held the dataset as saved, in ISO 8601, which load

accepts str

the same time for the same content.

load async

load(
    name: str,
    /,
    *,
    input_type: type[InputT],
    verdict_type: type[VerdictT],
    revision: str | None = None,
) -> Dataset[InputT, VerdictT]

Read the dataset's active items, as they are or as they were at revision.

Items not written by evalr are read as they are: the input as the input, and the expected output as the verdict.

Raises:

Type Description
DatasetNotFound

Langfuse has no dataset of that name, or the revision is not a time.

ValidationError

The items are not examples of these types.

Scores

LangfuseScoreSink

LangfuseScoreSink(client: Langfuse)

Records scores in Langfuse, on the traces or sessions they judge.

It records evaluators' scores and the libraries' mirrors of people's feedback alike. Each score keeps its id (Langfuse's score_id), so recording a score again replaces it, and its span (Langfuse's observation) and timestamp, when it has them. A yes or no is sent as 1 or 0, and the score's metadata (its source, the evaluator, its version and the confidence) becomes Langfuse's. Langfuse sends scores in the background; flush waits for them.

Use a Langfuse client.

Parameters:

Name Type Description Default
client Langfuse

The application's client.

required

record async

record(scores: Sequence[Score]) -> None

Queue the scores for Langfuse.

Raises:

Type Description
ValueError

A score has neither a trace nor a session, one of which Langfuse needs to attach it to; none of the scores is queued.

flush async

flush() -> None

Wait until Langfuse has sent everything queued.

LangfuseScoreConfigStore

LangfuseScoreConfigStore(client: Langfuse)

Keeps score configs in Langfuse, so it knows each score's type, range and choices.

Langfuse's API is synchronous, so its calls run in a worker thread.

Use a Langfuse client.

Parameters:

Name Type Description Default
client Langfuse

The application's client.

required

names async

names() -> set[str]

Return the names of Langfuse's score configs, archived or not.

create async

create(config: ScoreConfig) -> None

Create a score config in Langfuse.

Raises:

Type Description
ValueError

The name is not one Langfuse accepts: at most 35 letters, digits, spaces and _.()-. Shorten the type's name or the field's.

Experiments

LangfuseExperimentTracker

LangfuseExperimentTracker(
    client: Langfuse,
    *,
    type_names: Mapping[type[BaseModel], str] | None = None,
)

Runs experiments in Langfuse, where they show beside the traces of each item.

Each example's task and evaluators run in the item's trace, so the spans of whatever they call nest under it, and each verdict records it. Text fields, which Langfuse's evaluations cannot hold, become the comment of the verdict's other scores.

Use a Langfuse client.

Parameters:

Name Type Description Default
client Langfuse

The application's client.

required
type_names Mapping[type[BaseModel], str] | None

Score names for verdict types, where a library registers its feedback under a name other than the class name in snake case, so that evaluators' scores and people's line up.

None

run_experiment async

run_experiment(
    name: str,
    /,
    *,
    dataset: Dataset[InputT, VerdictT],
    task: Task[InputT, VerdictT, OutputT],
    evaluators: Sequence[Evaluator[OutputT, BaseModel]],
    run_name: str | None = None,
    max_concurrency: int = 4,
    metadata: Mapping[str, str] | None = None,
) -> ExperimentResult[OutputT]

Run the task on every example in Langfuse, and judge each output.

A run is named by Langfuse ({name} - {time}) unless named.