The Open Source AI Stack: A Practitioner's Guide to Building Agentic Systems — Spring 2026

Every quarter, someone publishes a definitive guide to the AI tech stack. By the time you finish reading it, half the recommendations are outdated. One framework has gone to maintenance mode. Another has shipped a breaking rewrite. A third has been absorbed into a hyperscaler's proprietary offering.

This is not that guide.

This is a seasonal snapshot — our Spring 2026 edition of the open source AI stack for building agentic systems in production. We plan to refresh it in summer, fall, and winter, because the best architecture decision you can make in 2026 is one you can revisit in 2027. If you're reading this after June 2026, check for the newer edition. The landscape moves that fast.

At Facet Interactive, we've watched open-source ecosystems mature three times — Drupal, React, dbt. The agentic AI stack is following the same curve, just compressed into months instead of years. In over a decade of building on open source across industries — fromlaw firms tomedical spas to professional services — we've learned one thing the hard way: the teams that choose architectures they can evolve outperform the teams that chase the "best" tool at any single point in time. That principle — what we call theFour Foundations: analytics, tooling, process, and automation — is how we systemize growth for our clients, and it applies to the open source AI stack as much as it does to any digital foundation.

A note on shelf life: This guide reflects the open source AI stack as of March 2026. Some recommendations will age well. Others won't survive the summer. We've flagged the areas of highest volatility in each section. If a specific tool recommendation feels wrong when you read this, it probably changed — and that's exactly why we version these guides.

The Eight-Layer Architecture

Building a production-grade agentic system isn't a single technology choice — it's eight. Each layer serves a distinct purpose, and getting one wrong creates problems that ripple through the rest. This architecture is informed by theResponsible AI Labs (RAIL) framework and refined through our own production deployments.

Layer Purpose Our Spring 2026 Pick Volatility

1. Languages

Core development

Python (primary), TypeScript (frontends)

Low

2. ML Frameworks

Model training & inference

PyTorch, HuggingFace Transformers

Low

3. LLM Providers

Model access

Anthropic Claude + open-source Llama/Mistral

Medium

4. Orchestration

Agent workflows & state

LangGraph

Medium-High

5. Document Ingestion

Static assets → RAG-ready chunks

Docling (IBM Research)

Medium

6. Vector & Data Storage

Semantic memory + structured data

PostgreSQL + pgvector

Low

7. Observability

LLM tracing + infra monitoring + analytics

Langfuse + SigNoz + PostHog

Medium

8. Safety & Evaluation

Responsible AI guardrails

RAIL Score API + Guardrails AI

Medium-High

This isn't an exhaustive catalog. It's an opinionated stack — what we'd deploy today for a team building agentic systems on open source LLM infrastructure. Every pick has trade-offs, and we'll walk through them.

Orchestration: LangGraph Is the Current Front-Runner

The orchestration layer is where your open source AI stack lives or dies. This is the layer that determines how agents reason, plan, retry, and coordinate — and it's also the layer with the highest turnover risk.

Why LangGraph (for now): LangGraph gives you state graphs where each business process is a graph, nodes are agents or tools, and edges are conditions determining which node fires next. It supports persistent state (backed by Postgres or Redis), loops for iterative refinement, and human-in-the-loop checkpoints. The LangChain ecosystem behind it has128,000+ GitHub stars and according to theirState of Agent Engineering report, 57% of surveyed respondents now have agents in production.

The competition and why it matters: CrewAI remains the fastest path to a prototype — its role-based agent model maps intuitively to how teams think about delegation. AutoGen pioneered conversational multi-agent patterns. But here's the velocity problem in action: Microsoftshifted AutoGen to maintenance mode in favor of their broader Agent Framework. If your production system was built on AutoGen, you're now planning a migration.

Framework Best For State Mgmt Production-Ready Risk

LangGraph

Complex stateful workflows

Excellent

Yes

Vendor coupling to LangChain ecosystem

CrewAI

Rapid prototyping, role-based teams

Good

Growing

Smaller community, fewer battle-scars

AutoGen

Conversational multi-agent

Good

Maintenance mode

Microsoft shifted focus

Pydantic AI

Type-safe agent development

Good

Emerging

Early ecosystem

Action: If you're starting a new agentic project today, default to LangGraph. If you're already on CrewAI and it's working, don't migrate for the sake of migrating. If you're on AutoGen, start planning your exit.

Document Ingestion: The Layer Most Teams Get Wrong

RAG pipelines fail at ingestion, not retrieval. If your documents go in as poorly parsed text blobs, no amount of embedding optimization will save your retrieval quality.

Docling is the open-source answer from IBM Research — a full document conversion toolkit that understands layout, reading order, tables, code blocks, and formulas. It handles PDFs, DOCX, PPTX, HTML, images, and more, converting everything into structured representations with semantic chunking that respects document hierarchy. The recent release ofGranite-Docling — a 258M parameter vision-language model — enables end-to-end document parsing in a single model pass, dramatically simplifying the pipeline.

Your FOSS AI development ingestion pipeline becomes:

  1. Raw documents land in object storage (S3/MinIO)
  2. Docling converts them to structured DoclingDocument representations
  3. Docling's chunker produces semantically coherent chunks with metadata
  4. Chunks + embeddings are written to PostgreSQL + pgvector

Compared to generic text extractors, Docling preserves document structure and semantics — which is the difference between an agent that can actually interpret your contracts and one that hallucinates section references.

Vector and Data Storage: PostgreSQL + pgvector Is the Pragmatic Choice

You don't need a dedicated vector database if you already run PostgreSQL. The pgvector extension gives you vector similarity search alongside your existing relational data — no new infrastructure to manage, no new backup strategy, no new access control layer.

For most teams building an agentic AI tech stack at SMB scale (50-500 employees, the organizations we work with most), pgvector handles the load. You add a dedicated vector database (Qdrant, Weaviate, Pinecone) when you hit scale problems — millions of vectors, sub-10ms latency requirements, or multi-tenant isolation needs that justify the operational overhead.

Database Best For Cost Operational Complexity

pgvector

Teams already on PostgreSQL, moderate scale

Low

Low (it's just Postgres)

Qdrant

High-performance retrieval, rich filtering

Low-Medium

Medium

Pinecone

Managed, enterprise-scale

Medium-High

Low (managed service)

Chroma

Prototyping, local development

Free

Low

Action: Start with pgvector. Migrate to a dedicated vector store only when you have data proving you need one — not because a blog post told you to.

Observability: Three Tools, Three Lenses

Agentic systems need observability at three levels, and no single tool covers all of them. This is where most open source LLM infrastructure guides stop at use Langfuse: and call it done.

Langfuse (LLM-native): Captures traces of every LangChain/LangGraph run — prompts, responses, tool calls, token usage, errors. Stores prompt versions and attaches evaluation scores. This is your "why did the agent do that?" lens.

SigNoz (infrastructure): OpenTelemetry-native observability for everything around the agents — HTTP latency, database queries, Model Context Protocol (MCP) gateway performance, container resources. This is your SRE lens.

PostHog (product analytics): User-facing metrics — which agent workflows get invoked, where users abandon, which cohorts get the most value. This is your "is it actually helping?" lens.

Each tool is open source and self-hostable. Together they give you the full picture: agent reasoning (Langfuse), system health (SigNoz), and business impact (PostHog).

Safety and Evaluation: The Layer Everyone Skips

Most open source AI stack guides treat safety as a footnote. We think it's a first-class architectural layer — and theResponsible AI Labs (RAIL) framework gives us a structured way to make it concrete.

The RAIL Score API evaluates AI outputs across eight dimensions: Fairness, Safety, Reliability, Transparency, Privacy, Accountability, Inclusivity, and User Impact. Integrating it into your pipeline means every agent output gets scored before it reaches a user — not as a manual review step, but as an automated guardrail.

For input and output guardrails — prompt injection detection, PII filtering, topic boundaries, content safety, factuality checks, and format validation —Guardrails AI provides an open-source framework with a hub of pre-built validators. Pair RAIL's evaluation scoring with Guardrails AI's runtime enforcement, and you have a responsible AI layer that's auditable and repeatable.

Action: Don't ship agents to production without at least basic input/output guardrails. The reputational cost of one unfiltered agent response far exceeds the engineering cost of adding safety checks.

Infrastructure: Portainer Over Docker, Not Kubernetes (Yet)

For teams that don't have a dedicated platform engineering group — which describes most of the 50-500 employee organizations we work with — a full Kubernetes deployment is overkill. Portainer over Docker Compose gives you per-stack lifecycle management, logs, configuration, secrets injection, and rollbacks without the operational overhead of a K8s cluster.

Deploy as separate Docker stacks: core app (LangGraph + LangChain + model gateway + Redis + Postgres), ingestion workers (Docling), observability (Langfuse + SigNoz + PostHog), and you get a platform without needing a platform team. This is the automation foundation of ourFour Foundations framework — infrastructure that systemizes deployment without requiring specialized platform knowledge.

Graduate to Kubernetes when your agent fleet grows beyond what a single node can handle, or when you need auto-scaling for variable inference workloads.

The Decision Framework: Budget, Team, and Timeline

Not every team needs every layer. Here's how to right-size your agentic AI tech stack based on budget and team capacity — the same kind of FOSS AI development right-sizing we apply across our Four Foundations work:

Budget Stack Recommendation Team Size

Under $1K/month

Commercial LLM APIs + LangGraph + pgvector + Langfuse (self-hosted)

1-2 engineers

$1K-$10K/month

Mix commercial + open-source models, Docling pipeline, full observability stack

2-4 engineers

Over $10K/month

Self-hosted models, GPU inference (vLLM), full MLOps, RAIL evaluation pipeline

4-8 engineers

The key insight: you don't have to build all eight layers on day one. Start with orchestration (LangGraph) and an LLM provider (commercial API). Add RAG when you need domain-specific knowledge. Add observability when you need to debug agent decisions. Add the safety layer before you ship to production users.

What Will Change by Summer 2026

We're flagging three areas where we expect the Spring 2026 picks to shift:

  1. Orchestration consolidation. The framework wars will produce winners and casualties. At least one more major framework will either get acquired or go to maintenance mode before Q3.
  2. Inference costs will drop again. Open-source model quality continues to close the gap with commercial APIs. By summer, the budget threshold for self-hosting may shift downward significantly.
  3. MCP standardization. The Model Context Protocol is becoming the standard for tool integration, but the spec is still evolving. Expect breaking changes and ecosystem churn.

This is why we version these guides. The open source AI stack isn't a set-and-forget decision — it's a quarterly architecture review.

Common Questions About Building an Open Source AI Stack

Do I need to go fully open source to build agentic systems?

No. Most production stacks are hybrid — commercial LLM APIs (OpenAI, Anthropic) paired with open-source orchestration, storage, and observability. The open source layers give you control where it matters most: your data, your workflows, and your ability to switch providers without rewriting your business logic.

How do I choose between LangGraph, CrewAI, and AutoGen?

LangGraph for production systems requiring sophisticated state management and conditional workflows. CrewAI for rapid prototyping with role-based agent teams. AutoGen is now in maintenance mode — we'd avoid it for new projects. The real question is whether your orchestration logic can survive a framework migration, which argues for keeping business logic separate from framework-specific code.

What's the minimum viable agentic stack?

LangGraph + a commercial LLM API + PostgreSQL. That's enough to build a stateful agent that can reason, use tools, and persist memory. Add layers as your requirements grow — but resist the urge to over-architect on day one.

How do I evaluate if my open source AI stack is production-ready?

Four criteria: Can you trace every agent decision? Can you roll back a bad deployment? Can you explain to a non-technical stakeholder why the agent made a specific choice? Can you swap out any single component without rewriting the rest? If the answer to any of these is "no," you have architecture work to do.

The Stack Is a Living Document — You Must Be Ready to Adapt

We've been building on open source for over a decade across industries — from Drupal sites for law firms to data pipelines for medspas to agent-driven workflow automation for professional services firms. The pattern is always the same: the teams that systemize their technology stack as a living architecture, reviewed and evolved quarterly, outperform the teams that "pick a stack" once and defend it for years.

The open source AI stack in Spring 2026 is LangGraph + Docling + PostgreSQL/pgvector + Langfuse + SigNoz + PostHog + Portainer + RAIL + Guardrails AI. By fall, some of these picks will change. The architecture principles won't.

If your team is evaluating an open source AI stack for agentic systems and wants a second opinion from a team that's been through three open-source ecosystem maturations,reach out for a stack review. We bring the sameAI strategy consulting approach we use across our client base — we'll tell you what we'd pick for your specific scale, budget, and timeline, and what we'd plan to revisit next quarter.

This is the Spring 2026 edition. Summer 2026 edition coming in July.