vigil

LLM observability · evaluation

See what your LLM application is actually doing.

Vigil records every model call your application makes — prompt, response, latency, model and tokens — and scores each response for relevance, so you know which answers need a closer look.

app.py
from vigil import Vigil

vigil = Vigil(service_name="support-bot")

with vigil.start_span("answer", span_type="llm") as span:
    span.set_input(question)
    answer = llm.complete(question)
    span.set_output(answer)
    span.record_llm_usage(provider="openai", model="gpt-4o-mini")

Questions Vigil answers

What is my application sending to the model?
Every prompt and response, with the trace of retrieval and tool calls around it.
Which responses missed the point?
Each LLM response gets a relevance score and a pass/fail label against its input.
Where does the time go?
Per-span latency in a waterfall, plus p50/p95 latency and error rate over time.
Is quality drifting?
Evaluation results and token usage by model, environment and release.

How it works

  1. 01

    Instrument

    Wrap model calls in spans with the Python SDK, or POST spans to the HTTP API from any language.

  2. 02

    Capture

    Vigil stores each trace — input, output, model, tokens, latency, errors — scoped to your project.

  3. 03

    Evaluate

    Enabled evaluators score LLM responses in the background on Vigil's own workers. Prompts never go to a third-party model.

From sign-up to your first evaluated trace in a few minutes.

Today: Python SDK and HTTP API, relevance evaluation. Groundedness and faithfulness evaluators are planned, not shipped.

Create an account