LLM observability · evaluation
See what your LLM application is actually doing.
Vigil records every model call your application makes — prompt, response, latency, model and tokens — and scores each response for relevance, so you know which answers need a closer look.
app.py
from vigil import Vigil
vigil = Vigil(service_name="support-bot")
with vigil.start_span("answer", span_type="llm") as span:
span.set_input(question)
answer = llm.complete(question)
span.set_output(answer)
span.record_llm_usage(provider="openai", model="gpt-4o-mini")Questions Vigil answers
- What is my application sending to the model?
- Every prompt and response, with the trace of retrieval and tool calls around it.
- Which responses missed the point?
- Each LLM response gets a relevance score and a pass/fail label against its input.
- Where does the time go?
- Per-span latency in a waterfall, plus p50/p95 latency and error rate over time.
- Is quality drifting?
- Evaluation results and token usage by model, environment and release.
How it works
- 01
Instrument
Wrap model calls in spans with the Python SDK, or POST spans to the HTTP API from any language.
- 02
Capture
Vigil stores each trace — input, output, model, tokens, latency, errors — scoped to your project.
- 03
Evaluate
Enabled evaluators score LLM responses in the background on Vigil's own workers. Prompts never go to a third-party model.
From sign-up to your first evaluated trace in a few minutes.
Today: Python SDK and HTTP API, relevance evaluation. Groundedness and faithfulness evaluators are planned, not shipped.