Build accountable AI.
Keep it under your control.
Langfuse is the open-source platform for observing, evaluating, and improving AI agents and LLM applications. Deploy across air-gapped, on-premises, and cloud environments— and keep sensitive prompts, outputs, traces, and evaluation data inside your approved infrastructure.
Understand every agent. Improve every outcome.
Government AI must do more than work in a demo. Teams need to understand how an agent reached an answer, measure whether it is reliable, and improve it without losing control of sensitive data.
Trace every model call, tool invocation, retrieval step, and agent decision. Investigate failures with the full context of each request, session, model, prompt, latency, and cost.
Observability docs ↗Score outputs with LLM-as-a-judge, deterministic checks, human review, and user feedback. Turn production failures into datasets and regression tests before the next release.
Evaluation docs ↗Monitor quality, security scores, latency, and cost. Set thresholds and route alerts through webhooks, Slack, or GitHub Actions — so teams can act before isolated failures become systemic.
Alerts docs ↗Mission-ready deployment, on your terms.
Langfuse is built to run where government teams already operate — behind a firewall, in a classified network, or in an approved cloud account — without changing the product or the data model.
Deploy Langfuse in a VPC, on premises, or in a fully air-gapped Kubernetes environment. Internet access is optional. Bring your own infrastructure, networking, storage, and operational controls.
Networking docs ↗The complete Langfuse repository is public. All core product capabilities — tracing, evaluations, prompt management, experiments, and annotation — are MIT-licensed without usage limits. Enterprise extensions live in clearly marked directories and activate only with a license key.
Open-source licensing ↗Self-hosted Langfuse is not a reduced fork. It uses the same codebase and architecture as Langfuse Cloud. Asynchronous ingestion absorbs traffic spikes, incoming events are persisted before processing, and background migrations reduce disruption during upgrades.
Architecture overview ↗Application traces create a detailed record of model calls and agent actions. Enterprise audit logs add immutable records of who changed what, when, and with which before-and-after state. SSO, role-based access control, SCIM, retention policies, and server-side data masking support centralized governance.
Audit logs ↗Security without giving up developer velocity.
Self-host Langfuse so application teams can debug and evaluate agents quickly, while security teams keep telemetry, prompts, and evaluation data inside the approved boundary.
For deployments with FIPS requirements, compliant Langfuse Docker images are available upon request.
Book a meeting↗Run the platform and its open-source dependencies in infrastructure you control.
Redact data in the SDK before transmission, or apply centralized ingestion masking in self-hosted Enterprise deployments.
Instrument with OpenTelemetry, or use Langfuse SDKs and integrations across models, frameworks, and languages.
Use versioned releases and deploy changes on your schedule.
Start locally. Deploy for the mission.
Run Langfuse locally with Docker Compose in minutes. The repository is public, the core product is MIT-licensed, and the same platform is already used in production at scale. Move to Kubernetes or Terraform without changing the product or data model.
$ git clone https://github.com/langfuse/langfuse.git $ cd langfuse $ docker compose up
Bring accountable AI into your environment.
See how Langfuse can help your team observe, evaluate, and improve mission-critical AI systems — without moving sensitive data outside your control.