ESC
Platform
Guard\ \ Block AI hallucinations in real-time with guardrails Evaluate\ \ Run comprehensive evaluations with 20+ metrics Error Feed\ \ Sentry-style error tracking for AI agents Simulations\ \ Simulate thousands of multi-turn conversations Scenarios\ \ Define branching conversation test scenarios Synthetic Data\ \ Generate diverse, realistic test data AI Optimization\ \ Continuous improvement with reinforcement learning Tracing\ \ End-to-end request tracing for AI agents Dashboards\ \ Custom dashboards with drag-and-drop widgets Alerting\ \ AI-powered alerts for anomalies and hallucination spikes Guardrails (Monitor)\ \ Real-time guardrail monitoring and block rate insights Datasets\ \ Manage and version evaluation datasets Experiments\ \ Structured experiments across models and prompts Agent IDE\ \ Build & test AI agents visually
Pages
Home\ \ Future AGI - AI agent hallucination detection platform Pricing\ \ Simple, transparent pricing. Start free, scale as you grow. Enterprise\ \ Enterprise-grade AI safety at scale Startups\ \ $10K in free credits and 6 months Pro access Roadmap\ \ Public product roadmap - see what we're building next Blog\ \ Guides, engineering deep-dives, and product updates Research\ \ Papers on hallucination detection, evaluation, and guardrails Customers\ \ Case studies from teams using Future AGI eBooks\ \ In-depth guides on AI agent evaluation and RAG Handbook\ \ The Flight Manual - how we work, what we believe
Docs
Introduction\ \ Future AGI is an AI lifecycle platform designed to support enterprises throughout their AI journey. It combines rapid prototyping, rigorous evaluation, continuous observability, and reliable deployment to help build, monitor, optimize, and secure generative AI applications. Self-Hosting\ \ Deploy the full Future AGI platform on your own infrastructure with Docker Compose or Kubernetes. Quickstart\ \ Future AGI is an AI lifecycle platform designed to support enterprises throughout their AI journey. It combines rapid prototyping, rigorous evaluation, continuous observability, and reliable deployment to help build, monitor, optimize, and secure generative AI applications. Setup Observability\ \ Set up Future AGI Observe for production monitoring. Configure auto-instrumented tracing for OpenAI, Anthropic, LangChain, and other LLM frameworks. Running Evals in Simulation\ \ Run evaluations in Future AGI simulations. Test AI agents against simulated customers and score interactions for quality, context retention, and escalation. Generate Synthetic Data\ \ Generate synthetic datasets with Future AGI. Define schemas, column types, and constraints to create realistic data for training and evaluation. Create Prompts\ \ Create and manage AI prompts in Future AGI's Prompt Workbench. Design, test, version, and optimize prompts with built-in model selection and evaluation. Setup MCP Server\ \ Set up the Future AGI MCP Server to interact with the platform via natural language from Claude, Cursor, or VS Code using Model Context Protocol. Annotations Quickstart\ \ Get started with annotations in 5 minutes -- create a label, set up a queue, add items, and start annotating. Prism AI Gateway Quickstart\ \ Make your first LLM request through Prism in under 5 minutes Overview\ \ Add human feedback to your AI outputs with annotation labels, queues, and scores across traces, datasets, prototypes, and simulations. Scores\ \ Understand the Score model -- the unified annotation primitive that stores labels, values, and metadata across all source types. Labels\ \ Create, configure, and manage annotation labels. Understand the five label types and when to use each. Queues\ \ Create and manage annotation queues: assignment strategies, multi-annotator support, review workflows, and queue lifecycle. Add Items to Queues\ \ Learn how to add traces, spans, sessions, dataset rows, prototypes, and simulation calls to annotation queues. Annotate Items\ \ Complete guide to the annotation workspace -- label inputs, keyboard shortcuts, navigation, instructions, and completion workflow. Inline Annotations\ \ Annotate traces, spans, sessions, and prototypes directly from their detail views without using queues. Analytics & Agreement\ \ Track annotation progress, annotator performance, label distribution, and inter-annotator agreement metrics. Export Annotations\ \ Export completed annotations as datasets (JSON/CSV) for fine-tuning, evaluation, or analysis. Automation Rules\ \ Set up rules to automatically add items to queues or pre-fill annotations based on conditions. Python SDK\ \ Annotate traces and manage annotation queues programmatically using the FutureAGI Python SDK. JavaScript SDK\ \ Annotate traces and manage annotation queues programmatically using the FutureAGI JavaScript/TypeScript SDK. Annotation Queue Using SDK\ \ Create and manage annotation queues programmatically using the Future AGI Python SDK. Overview\ \ Create, manage and analyze datasets for AI model development and evaluation Understanding Datasets\ \ How datasets work in Future AGI: structure, column types, creation methods, and lifecycle. Static Columns\ \ Static columns store fixed values in a dataset that only change when manually updated. Dynamic Columns\ \ Columns that are generated automatically by running prompts, models, or code against your dataset rows. Synthetic Data\ \ Generate realistic datasets from a schema without using real user data. Create New Dataset\ \ Learn to create datasets to do experimentations on them Add Rows to Dataset\ \ Learn how to add rows to your dataset Add Columns to Dataset\ \ Add static columns for fixed values or dynamic columns whose values are computed from other columns or external operations. Run Prompt in Dataset\ \ Learn how to execute prompts against your dataset and generate responses Experiments in Dataset\ \ To test, validate, and compare different prompt configurations Add Annotation\ \ Annotations are essential for refining datasets, evaluating model outputs, and improving the quality of AI-generated responses. Overview\ \ Automatically detect, cluster, and fix errors in your AI agent traces with Error Feed. Error Taxonomy\ \ Categories, subcategories, and descriptions of all error types detected by Error Feed. Using Error Feed\ \ How to read scores, insights, clusters, and recommendations from Error Feed. Overview\ \ Measure and compare quality of prompts and agents across datasets, simulations, and experiments. Understanding Evaluation\ \ How evaluation works in Future AGI: templates, judge models, results, and where evals run. Eval Types\ \ The four evaluation methods in Future AGI: LLM as Judge, Deterministic, Statistical Metric, and LLM as Ranker, and how modality affects which ones apply. Eval Templates\ \ What eval templates are, the difference between built-in and custom templates, and how output types work. Judge Models\ \ What a judge model is, how it scores responses, and how to choose the right one for your evaluation. Eval Results\ \ What eval results contain, how to read them, and how results are stored and aggregated across runs. Built-in Evals\ \ All built-in evaluation templates available on the platform. Evaluate via Platform & SDK\ \ Run evaluations via the Future AGI platform UI or the Python SDK. Create Custom Evals\ \ Define custom evaluation criteria and rules for your use case beyond built-in templates. Eval Groups\ \ Organize multiple evaluations into groups and run them together across datasets, simulations, and more. Use Custom Models\ \ Use your own or third-party models for evaluations via supported providers or a custom API endpoint. Future AGI Models\ \ Future AGI's proprietary models trained on a vast variety of datasets to perform evaluations. Evaluate CI/CD Pipeline\ \ Run Future AGI evaluations in your CI/CD pipeline to assess model performance on every pull request and keep quality checks consistent before deployment. Overview\ \ Store your organization’s content to ground synthetic data generation and evaluations in real source material. Understanding Knowledge Base\ \ What a Knowledge Base is, what content types are supported, and how files are processed. Create KB Using SDK\ \ Create and manage Knowledge Bases programmatically with the Future AGI Python SDK: create, update, add or remove files, and delete KBs from code or automation. Create KB Using UI\ \ Create and populate a Knowledge Base from the Future AGI platform: name it, upload documents, and wait for processing to finish. Overview\ \ Monitor and evaluate LLM applications in production with real-time tracing, session analysis, and alerting. Understanding Observability\ \ Core concepts behind LLM observability: what gets captured, how data is structured, and why it matters. What are Traces?\ \ In observability frameworks, a Trace is a comprehensive representation of the execution flow of a request within a system. It is composed of multiple spans, each capturing a specific operation or step in the process. Traces provide a holistic view of how different components interact and contribute to the overall behavior of the system. What are Spans?\ \ Understand spans in Future AGI tracing. Learn about span types including LLM, tool, chain, retriever, and embedding spans with their attributes. What is OpenTelemetry?\ \ Learn how Future AGI uses OpenTelemetry for vendor-neutral, high-performance tracing of AI applications with standardized telemetry collection. What is traceAI?\ \ Learn about traceAI, Future AGI's open-source package for standardized AI application tracing built on OpenTelemetry with framework-specific instrumentors. Set Up Observability\ \ Instrument your application and send traces to an Observe project so you can monitor LLM calls, latency, and cost in one place. Run Evals on Traces\ \ Run automated quality checks on your traced spans in Observe: filter spans, choose historic or continuous runs, set sampling and limits, and attach preset or custom evaluations. Sessions\ \ Group traces into sessions so you can view and analyze multi-turn conversations, chatbot flows, and per-session metrics in Observe. Users\ \ View all traces, sessions, and metrics per end user in one place so you can debug, analyze behavior, and optimize at the user level. Alerts & Monitors\ \ Define monitors on Observe project metrics (system or evaluation) and get notified by email or Slack when values cross a threshold. Voice Observability\ \ Connect a voice provider (Vapi, Retell) and get call logs as traces in Observe without any SDK instrumentation. Set Up Tracing\ \ Connect your application to Future AGI by registering a tracer provider and adding instrumentation with auto-instrumentors or manual OpenTelemetry spans. Instrument with traceAI Helpers\ \ Future AGI's traceAI library offers convenient abstractions to streamline your manual instrumentation process. Get Current Tracer and Span\ \ Access the active span or tracer at any point in your code to enrich it with additional attributes and context. Enriching Spans with Attributes, Metadata, and Tags\ \ Capture additional context beyond what standard frameworks provide by enriching your traces with custom attributes, metadata, tags, session IDs, user IDs, and prompt templates. Logging Prompt Templates & Variables\ \ Attach prompt template data to spans so Future AGI can surface it in the prompt playground for testing changes without deploying. Events, Exceptions, and Status\ \ OpenTelemetry (OTEL) provides support for adding Events, Exceptions, and Status into spans. Set Session ID and User ID\ \ Adding SessionID and UserID as attributes to Spans for Tracing Tool Spans Creation\ \ Manually trace tool functions alongside LLM calls by creating spans that capture inputs, outputs, and key events. Mask Span Attributes\ \ Redact sensitive inputs, outputs, images, and embeddings from spans before they are exported:using environment variables or TraceConfig in code. Advanced Tracing (OTEL)\ \ Explore manual context propagation, custom decorators, and sampling techniques for real-world async, multi-service, and high-volume tracing scenarios. FI Semantic Conventions\ \ Use standardized attribute keys for spans to ensure consistent, queryable trace data across LLM models, frameworks, and vendors. In-line Evaluations\ \ Run evaluations directly inside a traced span so results are automatically attached to that span in the Future AGI dashboard. Adding Annotations to your Spans\ \ Label spans with custom tags, human feedback, and notes using the bulk-annotation API. Langfuse Integration\ \ Integrate Future AGI evaluations with Langfuse to attach evaluation results directly to your Langfuse traces. Overview\ \ Auto-instrumentation for LLM applications across Python, JavaScript, and Java. OpenAI\ \ Set up auto-instrumentation for OpenAI with Future AGI tracing. Install traceAI-openai to capture chat completion, embedding, and tool call spans. Anthropic\ \ Set up auto-instrumentation for Anthropic Claude with Future AGI tracing. Install traceAI-anthropic to capture LLM spans, inputs, and outputs. AWS Bedrock\ \ Set up auto-instrumentation for AWS Bedrock with Future AGI tracing. Install traceAI-bedrock to capture model invocation spans and metadata. Vertex AI\ \ Set up auto-instrumentation for Vertex AI with Future AGI tracing. Install traceAI-vertexai to capture Gemini model invocation and response spans. Google GenAI\ \ Set up auto-instrumentation for Google GenAI with Future AGI tracing. Install traceAI-google-genai to capture Gemini model interaction spans. Google ADK\ \ Set up auto-instrumentation for Google ADK with Future AGI tracing. Install traceai-google-adk to capture agent and tool execution spans. Groq\ \ Set up auto-instrumentation for Groq with Future AGI tracing. Install traceAI-groq to capture high-speed inference spans and performance data. MistralAI\ \ Set up auto-instrumentation for Mistral AI with Future AGI tracing. Install traceAI-mistralai to capture model inference spans and metadata. Together AI\ \ Set up auto-instrumentation for Together AI with Future AGI tracing. Use traceAI-openai to capture inference spans from Together AI models. Ollama\ \ Set up auto-instrumentation for Ollama with Future AGI tracing. Use traceAI-openai to capture spans from Ollama's OpenAI-compatible local LLM API. Portkey\ \ Set up auto-instrumentation for Portkey with Future AGI tracing. Install traceAI-portkey to capture routed LLM call spans and gateway metrics. LangChain\ \ Set up auto-instrumentation for LangChain with Future AGI tracing. Install traceAI-langchain to capture chain, tool, and LLM call spans. LangGraph\ \ Set up auto-instrumentation for LangGraph with Future AGI tracing. Capture agent graph execution and state transition spans via LangChain instrumentor. LlamaIndex\ \ Set up auto-instrumentation for LlamaIndex with Future AGI tracing. Install traceAI-llamaindex to capture query, retrieval, and response spans. LlamaIndex Workflows\ \ Set up auto-instrumentation for LlamaIndex Workflows with Future AGI tracing. Trace workflow agent execution via the LlamaIndex instrumentor. LiteLLM\ \ Set up auto-instrumentation for LiteLLM with Future AGI tracing. Install traceAI-litellm to capture spans across multiple LLM provider calls. CrewAI\ \ Set up auto-instrumentation for CrewAI with Future AGI tracing. Install traceAI-crewai to capture crew task execution and agent interaction spans. AutoGen\ \ Set up auto-instrumentation for Autogen with Future AGI tracing. Install traceAI-autogen to capture multi-agent conversation spans automatically. Haystack\ \ Set up auto-instrumentation for Haystack with Future AGI tracing. Install traceAI-haystack to capture document pipeline and retrieval spans. DSPy\ \ Set up auto-instrumentation for DSPy with Future AGI tracing. Install traceAI-DSPy to capture program compilation and prediction spans automatically. OpenAI Agents\ \ Set up auto-instrumentation for OpenAI Agents SDK with Future AGI tracing. Install traceAI-openai-agents to capture agent workflow spans. Smol Agents\ \ Set up auto-instrumentation for Smol Agents with Future AGI tracing. Install traceAI-smolagents to capture lightweight agent execution spans. Instructor\ \ Set up auto-instrumentation for Instructor with Future AGI tracing. Install traceAI-instructor to capture structured output extraction spans. PromptFlow\ \ Set up auto-instrumentation for Prompt Flow with Future AGI tracing. Use traceAI-openai to capture prompt flow execution and LLM call spans. Guardrails\ \ Set up auto-instrumentation for Guardrails AI with Future AGI tracing. Install traceAI-guardrails to trace validation and LLM interaction spans. MCP\ \ Set up auto-instrumentation for MCP with Future AGI tracing. Install traceAI-mcp to capture Model Context Protocol server and tool call spans. Mastra\ \ Set up auto-instrumentation for Mastra with Future AGI tracing. Configure @traceai/mastra to export TypeScript agent spans to Future AGI. Vercel AI SDK\ \ Set up auto-instrumentation for Vercel AI SDK with Future AGI tracing. Install @traceai/vercel to capture AI function call spans in Next.js apps. LiveKit\ \ Integrate LiveKit with Future AGI for voice agent observability. Trace real-time voice interactions and monitor agent performance with traceAI-livekit. Pipecat\ \ Set up auto-instrumentation for Pipecat voice apps with Future AGI tracing. Install traceAI-pipecat to capture voice pipeline and processing spans. Overview\ \ Set up TraceAI for Java applications. Initialize the tracer, configure credentials, and instrument your LLM clients, vector databases, and frameworks. Spring Boot\ \ Add tracing to Spring Boot apps with Spring AI. Configure application.yml, wrap your ChatModel and EmbeddingModel, and traces are collected automatically. OpenAI\ \ Trace OpenAI chat completions, embeddings, and streaming responses in Java with TracedOpenAIClient. Anthropic\ \ Trace Anthropic Messages API calls in Java with TracedAnthropicClient. Uses reflection for cross-version compatibility. AWS Bedrock\ \ Trace AWS Bedrock model invocations in Java with TracedBedrockRuntimeClient. Supports both InvokeModel (raw JSON) and Converse (typed API). Cohere\ \ Trace Cohere chat, embedding, and reranking operations in Java with TracedCohereClient. Pinecone\ \ Trace Pinecone vector operations in Java with TracedPineconeIndex. Query, upsert, delete, and fetch with full span instrumentation. LLM Providers\ \ Trace Google GenAI, Vertex AI, Azure OpenAI, Ollama, and Watsonx in Java. All use the same Traced wrapper pattern. Vector Databases\ \ Trace vector database operations in Java. Qdrant, Milvus, ChromaDB, Weaviate, MongoDB, Redis, pgvector, Azure AI Search, and Elasticsearch. Frameworks\ \ Trace LangChain4j and Semantic Kernel operations in Java. Framework-level wrappers that instrument chains, agents, and prompt invocations. n8n\ \ With this integration, you can dynamically retrieve prompts from your Future AGI account, select specific versions, and compile prompts with variables - all within the familiar n8n interface. Overview\ \ Iteratively improve prompts using evaluation-driven feedback and optimization algorithms for higher-quality, more consistent AI responses. Understanding Optimization\ \ How prompt optimization works: the feedback loop, key components, algorithms, and how to choose the right one. Bayesian Search\ \ Use Bayesian optimization for few-shot prompt tuning: learns from trials to pick better example sets and configurations. Meta-Prompt\ \ A guide to the Meta-Prompt optimizer, which uses a teacher LLM for deep reasoning-based prompt refinement through systematic failure analysis and rewriting. ProTeGi\ \ A guide to ProTeGi (Prompt optimization with Textual Gradients), which systematically improves prompts by identifying failures, generating critiques, and applying targeted fixes. PromptWizard\ \ Learn about PromptWizard, a multi-stage feedback-driven optimizer that improves prompts through a cycle of mutation, critique, and refinement. GEPA\ \ Discover GEPA (Genetic Pareto), a powerful evolutionary algorithm that evolves prompts over generations using reflection and mutation for complex, high-stakes optimization. Random Search\ \ Understand the Random Search optimizer, a simple and effective gradient-free method for establishing a baseline in prompt optimization by exploring random variations. Using Python SDK\ \ Run prompt optimization from code with the agent-opt Python library. Using Platform\ \ Run prompt optimization from the Future AGI UI: pick a dataset and column, configure prompt and evals, run optimization, and apply the best prompt. Overview\ \ A unified API gateway for 100+ LLM providers with built-in guardrails, intelligent routing, caching, cost controls, and full observability. Core Concepts\ \ Understand the key building blocks of Prism: gateways, virtual API keys, organizations, providers, and configurations. API Reference\ \ Endpoints, request headers, and response headers for the Prism AI Gateway. Configuration\ \ How organization configuration works in Prism: sections, hierarchy, and real-time updates. Platform Integration\ \ How Prism AI Gateway connects to the broader Future AGI platform — observability, evaluation, protection, and experimentation. Manage Providers\ \ Add, configure, and manage LLM providers in Prism. Routing & Reliability\ \ Configure load balancing, failover, retries, and circuit breaking across LLM providers. Guardrails\ \ Set up safety guardrails to protect your LLM traffic with PII detection, prompt injection prevention, content moderation, and more. Caching\ \ Reduce costs and latency with Prism's exact match and semantic caching. Cost Tracking & Budgets\ \ Track LLM costs per request, set budget limits, and configure spend alerts. Streaming\ \ Use Server-Sent Events (SSE) streaming with Prism for real-time LLM responses. Shadow Experiments\ \ Mirror a percentage of production LLM traffic to alternative models for zero-risk evaluation. Rate Limiting\ \ Control request throughput to the Prism AI Gateway with configurable rate limits. MCP & A2A\ \ Connect AI agents to Prism using the Model Context Protocol (MCP) and Google's Agent-to-Agent (A2A) protocol. Self-Hosted\ \ Deploy Prism AI Gateway on your own infrastructure using Docker or a Go binary. Overview\ \ Create, manage, and optimize AI prompts for reliable and consistent language model outputs. Prompt Engineering\ \ What prompt engineering is, how to think about crafting effective prompts, and how the Prompt Workbench supports the iteration process. Understanding Prompts\ \ What a prompt is, how it is structured, how variables work, and how prompts connect to models in the Prompt Workbench. Versions and Labels\ \ How prompt versioning and deployment labels work in the Prompt Workbench. Create Prompt from Scratch\ \ Build a new prompt manually in the Prompt Workbench with full control over structure, model, parameters, and variables. Create from Existing Template\ \ Start from a pre-built prompt template in the Prompt Workbench and customize it for your use case. Create with AI\ \ Generate a new prompt from a plain-language description using the Generate with AI feature in the Prompt Workbench. Prompt Workbench Using SDK\ \ Create, version, and run prompt templates programmatically using the Future AGI SDK (TypeScript/JavaScript or Python). Linked Traces\ \ Associate prompts with production traces to monitor latency, token usage, and cost per prompt version in the Prompt Workbench. Manage Folders\ \ Organize prompt templates into folders in the Prompt Workbench to keep your workspace navigable as your library grows. Overview\ \ Future AGI's Protect module brings real-time safety and policy enforcement directly into your GenAI application flow. Use Cases\ \ Future AGI's Protect acts as a vital guardrail for AI applications, ensuring security, reliability, and ethical compliance during real-time interactions across text, image, and audio modalities. Run Protect via SDK\ \ Set up and configure Protect to apply real-time safety checks to your AI application's inputs and outputs. Overview\ \ Build, test, and run multi-step AI workflows visually, no code required. Connect prompts, models, and agents on a drag-and-drop canvas. Understanding Agent Playground\ \ Learn the core building blocks of Agent Playground: graphs, nodes, ports, edges, and node templates. Versions & Execution\ \ Understand the version lifecycle, execution model, data routing, and batch execution in Agent Playground. Create a Graph\ \ Create your first agent graph, manage metadata, and work with versions in Agent Playground. Build a Workflow\ \ Add nodes, configure them, and connect them into an AI agent pipeline using the visual graph editor. Run & Monitor\ \ Execute agent workflows, view real-time results per node, and inspect execution history. Overview\ \ Test and compare LLM configurations, prompts, and parameters before deploying to production. Understanding Prototype\ \ What Prototype is, the problem it solves, and how versions, traces, and evals work together before you ship. Versions and Runs\ \ What a version is in Prototype, how runs get tagged to a version, and how the dashboard uses versions to compare configurations. Set Up Prototype\ \ Configure your environment, register your prototype project, and instrument your app so traces and evals appear in the Prototype dashboard. Evals\ \ Define which evaluations run on your prototype outputs using EvalTags, mapping, and optional custom evals. Choose Winner\ \ Rank prototype versions by evaluation scores, cost, and latency, then select and promote the best-performing version to production. Admin & Settings\ \ Learn how to access and manage your Future AGI API keys and secret keys from the developer dashboard for authentication. API Keys\ \ Create and manage API keys for authenticating with Future AGI SDKs and APIs. Profile & Security\ \ Manage your profile information, password, two-factor authentication, and passkeys. Organization Settings\ \ Configure your organization name and security policies. User Management\ \ Invite users, assign roles, and manage team members across your organization. Workspace Management\ \ Create and configure workspaces to organize projects, teams, and resources. AI Providers\ \ Configure LLM providers and custom models for evaluations, optimization, and other platform features. Integrations\ \ Connect Future AGI to external tools for observability, alerting, analytics, and log archival. Usage Summary\ \ Track API calls, token usage, and evaluation runs across your organization and workspaces. Billing & Pricing\ \ Manage your subscription, add funds, configure auto-reload, and view invoices. Roles & Permissions\ \ Resources Installation\ \ Install the Future AGI SDK and configure it for your project. FAQ\ \ Find answers to common questions about Future AGI products. Release Notes\ \ Latest Future AGI release notes covering new features, improvements, and bug fixes across datasets, evaluations, simulation, and observability products. Overview\ \ Test AI agents and prompts through controlled simulations before deploying to production. Agent Definition\ \ An agent definition is a configuration that specifies how your AI agent behaves during voice or chat conversations Scenarios\ \ Scenarios defines the test cases, customer profiles, and conversation flows that your AI agent will encounter during simulations. Personas\ \ Create personas that represent the customers or users in your simulation tests for more realistic scenarios. Run Voice Simulation\ \ Create and run voice simulation tests from the platform to test your agent against scenarios. Chat Simulation Using SDK\ \ Run Future AGI chat simulations from Python by providing an agent callback and executing an existing Run Test. Replay\ \ Replay real production sessions in a dev environment using chat simulation to debug, iterate, and improve your agent. Prompt Simulation\ \ Test your prompts in realistic multi-turn conversations directly from the Prompt Workbench — no agent deployment or SDK required. Evaluate Tool Calling\ \ Evaluate the tool-calling capabilities of your agent in simulation runs. View Results\ \ Read simulation results: transcripts, evaluation scores, performance analytics, and call logs. Fix My Agent\ \ In-depth diagnostics and targeted fixes for your agent's performance issues based on simulation results Overview\ \ Connect Future AGI with your existing AI frameworks, LLM providers, and tools. OpenAI\ \ Integrate OpenAI with Future AGI for auto-instrumented tracing. Capture chat completions, embeddings, and tool calls with traceAI-openai. Anthropic\ \ Integrate Anthropic Claude with Future AGI for auto-instrumented tracing. Install traceAI-anthropic and capture LLM calls with full observability. AWS Bedrock\ \ Integrate AWS Bedrock with Future AGI for auto-instrumented tracing. Capture model invocations and monitor performance with traceAI-bedrock. Vertex AI\ \ Integrate Vertex AI (Gemini) with Future AGI observability. Trace model calls and monitor performance using traceAI-vertexai instrumentation. Google GenAI\ \ Integrate Google GenAI with Future AGI observability. Set up traceAI-google-genai to capture model calls and monitor performance automatically. Google ADK\ \ Integrate Google ADK with Future AGI for auto-instrumented tracing. Monitor Google AI agent calls and tool usage with traceAI-google-adk. Groq\ \ Integrate Groq with Future AGI observability. Set up traceAI-groq to automatically trace high-speed inference calls and monitor LLM performance. MistralAI\ \ Integrate Mistral AI with Future AGI observability. Set up traceAI-mistralai to capture model calls and monitor inference performance automatically. Together AI\ \ Integrate Together AI with Future AGI observability. Trace inference calls to Together AI models using the traceAI-openai compatible package. Ollama\ \ Integrate Ollama with Future AGI observability. Trace locally-hosted LLM calls using the traceAI-openai package with Ollama's OpenAI-compatible API. Portkey\ \ Integrate Portkey AI gateway with Future AGI observability. Trace routed LLM calls and monitor performance with traceAI-portkey instrumentation. LangChain\ \ Integrate LangChain with Future AGI for auto-instrumented tracing. Capture chain executions, tool calls, and LLM interactions with traceAI-langchain. LangGraph\ \ Integrate LangGraph with Future AGI observability. Trace agent graph execution, tool usage, and state transitions using the LangChain instrumentor. LlamaIndex\ \ Integrate LlamaIndex with Future AGI observability. Set up traceAI-llamaindex to trace queries, retrieval, and response generation automatically. LlamaIndex Workflows\ \ Integrate LlamaIndex Workflows with Future AGI. Trace workflow-based agent execution and data processing using the LlamaIndex instrumentor. LiteLLM\ \ Integrate LiteLLM with Future AGI observability. Set up traceAI-litellm to trace calls across multiple LLM providers through a unified interface. CrewAI\ \ Integrate CrewAI with Future AGI observability. Set up traceAI-crewai to trace multi-agent crew task execution and tool usage automatically. AutoGen\ \ Integrate Autogen with Future AGI observability. Set up traceAI-autogen for automatic tracing of multi-agent conversations and workflows. Haystack\ \ Integrate Haystack with Future AGI observability. Set up traceAI-haystack to trace document processing pipelines and LLM calls automatically. DSPy\ \ Integrate DSPy with Future AGI observability. Set up traceAI-DSPy to automatically trace DSPy program compilation and inference pipelines. OpenAI Agents\ \ Integrate OpenAI Agents SDK with Future AGI. Trace agent tool calls, handoffs, and reasoning steps automatically with traceAI-openai-agents. Smol Agents\ \ Integrate Smol Agents with Future AGI observability. Set up traceAI-smolagents to trace lightweight agent tool calls and reasoning automatically. Instructor\ \ Integrate Instructor with Future AGI observability. Trace structured LLM output extraction and validation automatically using traceAI-instructor. PromptFlow\ \ Integrate Prompt Flow with Future AGI observability. Trace prompt flow executions and LLM calls automatically using the traceAI-openai package. Guardrails\ \ Integrate Guardrails AI with Future AGI observability. Trace guardrail validations and LLM interactions automatically using traceAI-guardrails. MCP\ \ Integrate Model Context Protocol (MCP) with Future AGI. Trace MCP server interactions and tool calls with traceAI-mcp auto-instrumentation. Mastra\ \ Integrate Mastra with Future AGI for TypeScript agent observability. Configure trace export using the @traceai/mastra package for LLM monitoring. Vercel AI SDK\ \ Integrate Vercel AI SDK with Future AGI. Set up @traceai/vercel for automatic tracing of AI-powered Next.js and Vercel applications. LiveKit\ \ Integrations Pipecat\ \ Integrate Pipecat with Future AGI for voice application observability. Trace and monitor voice pipelines with OpenTelemetry-based traceAI-pipecat. Overview\ \ Set up TraceAI for Java applications. Initialize the tracer, configure credentials, and instrument your LLM clients, vector databases, and frameworks. Spring Boot\ \ Add tracing to Spring Boot apps with Spring AI. Configure application.yml, wrap your ChatModel and EmbeddingModel, and traces are collected automatically. OpenAI\ \ Trace OpenAI chat completions, embeddings, and streaming responses in Java with TracedOpenAIClient. Anthropic\ \ Trace Anthropic Messages API calls in Java with TracedAnthropicClient. Uses reflection for cross-version compatibility. AWS Bedrock\ \ Trace AWS Bedrock model invocations in Java with TracedBedrockRuntimeClient. Supports both InvokeModel (raw JSON) and Converse (typed API). Cohere\ \ Trace Cohere chat, embedding, and reranking operations in Java with TracedCohereClient. Pinecone\ \ Trace Pinecone vector operations in Java with TracedPineconeIndex. Query, upsert, delete, and fetch with full span instrumentation. LLM Providers\ \ Trace Google GenAI, Vertex AI, Azure OpenAI, Ollama, and Watsonx in Java. All use the same Traced wrapper pattern. Vector Databases\ \ Trace vector database operations in Java. Qdrant, Milvus, ChromaDB, Weaviate, MongoDB, Redis, pgvector, Azure AI Search, and Elasticsearch. Frameworks\ \ Trace LangChain4j and Semantic Kernel operations in Java. Framework-level wrappers that instrument chains, agents, and prompt invocations. n8n\ \ With this integration, you can dynamically retrieve prompts from your Future AGI account, select specific versions, and compile prompts with variables - all within the familiar n8n interface. Langfuse\ \ Pull your existing Langfuse traces, spans, and scores into Future AGI automatically. Datadog\ \ Forward Prism Gateway logs and metrics from Future AGI to Datadog automatically. PostHog\ \ Send LLM usage events from Future AGI's Prism Gateway to PostHog for product analytics. Mixpanel\ \ Send LLM usage events from Future AGI's Prism Gateway to Mixpanel for product analytics. PagerDuty\ \ Route Future AGI alerts to PagerDuty so your on-call team gets paged when something breaks. Cloud Storage\ \ Archive Prism Gateway logs to S3, Azure Blob Storage, or Google Cloud Storage as compressed JSONL files. Message Queues\ \ Stream Prism Gateway logs to Amazon SQS or Google Pub/Sub for real-time processing. Overview\ \ Practical guides and tutorials for using Future AGI products effectively Running Your First Eval\ \ Score LLM outputs for hallucination, toxicity, and custom quality criteria — from local metrics to LLM-as-Judge. Custom Eval Metrics: Write Your Own Evaluation Criteria\ \ Define quality criteria in plain English and run them as reusable eval metrics from the dashboard or SDK on any dataset or production trace. Hallucination Detection with Faithfulness & Groundedness\ \ Score RAG outputs for faithfulness and groundedness to catch hallucinations before they reach users. RAG Pipeline Evaluation: Debug Retrieval vs Generation\ \ Score retrieval quality and generation quality independently to pinpoint whether your RAG pipeline is failing at retrieval or generation. Multimodal Evaluation: Images, Audio, and PDF\ \ Score image captions, detect AI-generated images, evaluate audio quality and TTS accuracy, and verify OCR output against source PDFs using built-in eval metrics. Tone, Toxicity, and Bias Detection Evals\ \ Evaluate LLM outputs for professional tone, harmful content, and demographic bias using the evaluate() function in a customer service scenario. Evaluate Customer Agent Conversations\ \ Score multi-turn conversations for quality, context retention, query handling, loop detection, escalation, and prompt conformance using built-in Turing metrics. Dataset SDK: Upload, Evaluate, and Download Results\ \ Upload a CSV, run batch evaluations across every row, and download scored results: all from the SDK. Async Evaluations for Large-Scale Testing\ \ Fire-and-forget async evaluations, poll for results, and run parallel evals across hundreds of items using the Evaluator SDK. Text-to-SQL Evaluation\ \ Evaluate LLM-generated SQL queries using the built-in text_to_sql Turing metric, local string comparison, and execution-based validation against a live database. Chat Simulation: Run Multi-Persona Conversations via SDK\ \ Use FutureAGI's Chat Simulation feature to define personas, generate scenarios, execute multi-turn conversations via the SDK, and diagnose failures with Fix My Agent. Voice Simulation: Define Agents, Personas, and Run Call Tests\ \ Use Voice Simulation to define voice agents with provider credentials, build caller personas with accent and speed controls, generate call scenarios, run parallel call tests with evaluations, and diagnose failures with Fix My Agent. Tool-Calling Agent Simulation with Tracing\ \ Run a tool-calling agent through simulated scenarios, trace every tool invocation as child spans, and inspect results in the Tracing dashboard. Simulate from the Prompt Workbench\ \ Run a simulation against your prompt directly from the FutureAGI Prompts page — no SDK, no code required. Create and Manage Datasets from the Dashboard\ \ Create a dataset, add columns, enter rows manually, import from CSV, run evaluations, and export — all from the FutureAGI dashboard, no code required. Synthetic Data Generation: Create Test Datasets from a Schema\ \ Use FutureAGI's Synthetic Data Generation feature to define column schemas, set categorical distributions, and generate structured test datasets — no code required. Annotate Datasets with Human-in-the-Loop Workflows\ \ Create annotation views, define labels, assign annotators, and log annotations programmatically via the SDK. Import Datasets from Hugging Face\ \ Pull any public Hugging Face dataset into FutureAGI with a single SDK call and run evaluations on it. Dynamic Dataset Columns: Enrich Rows with AI-Generated Data\ \ Use Dynamic Columns to add AI-generated summaries, sentiment labels, extracted entities, vector-retrieved context, parsed JSON fields, and conditional routing to any dataset — no code required. Prompt Versioning: Create, Label, and Serve Prompt Versions\ \ Use FutureAGI's Prompt Versioning feature to create prompt templates, commit numbered versions, assign labels like production, and serve the right version at runtime via SDK. Prototype and Iterate on LLM Applications\ \ Register a prototype project, auto-evaluate spans with EvalTags, iterate with versioned prompts, compare versions, and choose the winner before deploying to production. Manual Tracing: Add Custom Spans to Any Application\ \ Instrument any Python application with custom spans, user context, and metadata - and see every call visualized in the FutureAGI Tracing dashboard. Session-Based Observability for Multi-Turn Conversations\ \ Group every span from a multi-turn chatbot by session and user ID so conversations appear as a single, filterable unit in the FutureAGI Tracing dashboard. Monitoring & Alerts: Track LLM Performance and Set Quality Thresholds\ \ Generate rich trace data from a multi-step RAG agent, analyze historical performance trends in the Charts tab, and configure alerts with thresholds and notifications. Inline Evals in Tracing: Score Every Response as It's Generated\ \ Attach quality scores directly to production traces so you can see faithfulness, toxicity, and custom evals alongside every LLM call in FutureAGI Tracing. Distributed Tracing: Connect Spans Across Services\ \ Propagate OpenTelemetry trace context across microservices so every span - from your API gateway to your LLM backend - shows up in a single trace. Prompt Optimization: Improve a Prompt Automatically\ \ Use the agent-opt SDK to take a weak baseline prompt, run automated optimization, and deploy the best-performing variant - no manual prompt engineering required. Compare Optimization Strategies: ProTeGi, GEPA, and PromptWizard\ \ Run three optimization algorithms on the same task with different evaluation metrics and compare results to pick the best strategy for your use case. Dataset Optimization: Improve Prompts Directly in Your Dataset\ \ Use the dashboard Optimization tab to run automated prompt improvement on any Run Prompt column: no SDK code required. Protect: Add Safety Guardrails to LLM Outputs\ \ Use FutureAGI Protect to screen text for prompt injection, PII, toxicity, and bias with a single API call — stack multiple safety rules and switch to Protect Flash for high-volume pipelines. Knowledge Base: Upload Documents and Query with the SDK\ \ Upload documents to a Knowledge Base, manage files programmatically with the SDK, and use Knowledge Bases for grounded evaluations and synthetic data generation. Experimentation: Compare Prompts and Models on a Dataset\ \ Use the Experimentation feature to run multiple prompt variants across different models on the same dataset, evaluate outputs, and pick the winning configuration. Evaluation-Driven Development: Score Every Prompt Change Before Shipping\ \ Build a local eval loop that scores prompts against a test suite, compare before-and-after results, and gate promotion on quality thresholds. CI/CD Eval Pipeline: Automate Quality Gates in GitHub Actions\ \ Set up FutureAGI's CI/CD Eval Pipeline to run automated quality gates on every pull request, failing builds when eval scores drop below your configured thresholds. Agent Compass: Surface Agent Failures Automatically\ \ Instrument your AI agent with tracing, let Agent Compass analyze traces for errors, and review clustered failure patterns with actionable recommendations in the Feed dashboard. Using FutureAGI Evals\ \ Use FutureAGI Evals to evaluate your AI models Using FutureAGI Protect\ \ Use FutureAGI Protect to protect your data Using FutureAGI Dataset\ \ Use FutureAGI Dataset to create and manage your datasets Using FutureAGI KB\ \ Use FutureAGI Knowledge Base to create and manage your knowledge base Portkey Integration\ \ Combine Portkey and Future AGI for end-to-end LLM observability. Benchmark multiple models on response quality, latency, and cost. LangChain/LangGraph\ \ Add observability and evaluation to LangChain and LangGraph agents using Future AGI's tracing SDK for completeness, groundedness, and hallucination detection. LlamaIndex PDF RAG\ \ Build a production-ready LlamaIndex PDF RAG chatbot with Future AGI observability, tracing, and real-time evaluation of retrieval quality. CrewAI Research Team\ \ Learn how to build a multi-agent research system using CrewAI with integrated observability and in-line evaluations from FutureAGI for real-time quality monitoring. MongoDB\ \ Learn how to build production-grade PDF RAG chatbots using MongoDB Atlas for vector search and Future AGI to trace, evaluate, and real-time performance monitoring of LLM pipelines Meeting Summarization\ \ Evaluate meeting summarization quality using Future AGI. Score AI-generated summaries from transcripts for accuracy and completeness. AI SDR Evaluation\ \ Evaluate AI-generated sales outreach messages using Future AGI. Score SDR openers for relevance, personalization, and value proposition alignment. AI Agents Evaluation\ \ Evaluate AI agent function-calling and response quality using Future AGI's evaluation SDK with metrics like tool use accuracy and safety. Image Evaluation\ \ Evaluate AI-generated images for description alignment, artistic requirements, and replacement quality using the Future AGI SDK. Implement Observability\ \ Master AI observability with FutureAGI. Track LLM performance, monitor metrics, and optimize Python apps. Step-by-step guide with examples. Text-to-SQL Evaluation\ \ Build and evaluate a Text-to-SQL agent with Future AGI. Test natural language to SQL conversion accuracy using automated evaluation metrics. RAG with LangChain\ \ Experiment with LangChain RAG configurations using Future AGI. Build and evaluate a retrieval-augmented generation app with OpenAI embeddings. Evaluate RAG Apps\ \ Evaluate RAG applications with Future AGI using context adherence, retrieval quality, answer correctness, and other retrieval-augmented generation metrics. Trustworthy RAG Chatbots\ \ Evaluate RAG chatbot trustworthiness across retrieval accuracy, prompt injection resilience, privacy compliance, and tone adaptation with Future AGI. Decrease RAG Hallucination\ \ Reduce hallucinations in RAG pipelines by benchmarking chunking, retrieval, and chain strategies with Future AGI's evaluation suite. End-to-End Prompt Optimization\ \ Optimize prompts end-to-end with Future AGI. Learn evaluation-driven prompt refinement using automated scoring and version tracking. Basic Prompt Optimization\ \ A hands-on guide to optimizing your first prompt using the agent-opt Python library with a simple Random Search strategy. GEPA Optimization\ \ A guide to using GEPA, a powerful evolutionary algorithm for state-of-the-art prompt optimization in complex, high-stakes scenarios. Eval Metrics for Optimization\ \ Learn how to use the FutureAGI platform, local LLM-as-a-judge, and local heuristic metrics to guide your prompt optimization. Compare Strategies\ \ A practical guide to selecting the best optimization strategy (Bayesian Search, Meta-Prompt, GEPA, etc.) based on your specific task and goals. Import Datasets\ \ Learn how to prepare and integrate datasets from various sources (in-memory, CSV, JSON, JSONL) for effective prompt optimization. Chat Simulation with Fix My Agent\ \ Simulate AI chat agents at scale and get instant AI-powered diagnostics to improve performance Simulate SDK Demo\ \ This cookbook demonstrates how to use the agent-simulate SDK to test a conversational voice AI agent. Error Feed with Google ADK\ \ Set up a multi-agent system using Google ADK, send traces to Future AGI, and analyze agent errors with Error Feed. SDK Overview\ \ Evaluate LLM outputs, trace AI calls, optimize prompts, and test voice agents. Python, TypeScript, Java, and C# supported. Overview\ \ Evaluate LLM outputs with 76+ local metrics, cloud Turing models, or custom LLM-as-Judge criteria. Part of the ai-evaluation Python package. Running Evaluations\ \ Run evaluations with the evaluate() function — local heuristics, cloud Turing, or LLM-as-Judge, auto-routed based on your inputs. Distributed Evaluator\ \ Run evaluations at scale with blocking, async, or distributed execution. Backends for ThreadPool, Celery, Ray, Temporal, and Kubernetes. Built-in resilience. AutoEval\ \ Auto-generate evaluation pipelines from app descriptions. Pre-built templates for customer support, RAG, code assistants, healthcare, and more. Guardrails\ \ Screen AI inputs and outputs with model-based safety checks and fast local scanners. 14 guard models, 14 scanners, async and batch support. Local & Hybrid\ \ Run evaluations locally with zero API calls. Auto-route between local and cloud metrics. Use Ollama for offline LLM-based scoring. OpenTelemetry\ \ Built-in OpenTelemetry for the AI evaluation SDK. Auto-instrument LLM calls, track costs, enrich spans with scores, and export to any backend. Code Security\ \ AST-based vulnerability detection for AI-generated code. 15 detectors, 4 evaluation modes, multi-language support, built-in benchmarks, and dual-judge scoring. Overview\ \ Browse all 76+ local evaluation metrics by category. String checks, JSON validation, similarity, hallucination, RAG, agents, structured output, and guardrails. String & Similarity\ \ 23 local metrics for keyword matching, regex, length checks, BLEU, ROUGE, Levenshtein, and embedding similarity. JSON & Structured\ \ 14 metrics for validating JSON correctness, schema compliance, type checking, and structured output quality. Hallucination\ \ Detect hallucinations, unsupported claims, and contradictions in LLM outputs. 5 context-grounded metrics with optional NLI and LLM augmentation. RAG\ \ 19 local metrics for evaluating RAG pipelines — retrieval quality, generation faithfulness, advanced reasoning, and composite scores. Agents & Functions\ \ 11 metrics for evaluating agent trajectories, tool use, reasoning quality, and function call correctness. All run locally via evaluate(). Guardrails\ \ Security-focused scanner metrics that detect prompt injection, PII, secrets, and SQL injection in under 10ms. Cloud Evals\ \ Run pre-built evaluation templates on Future AGI's Turing cloud models. 100+ templates covering safety, RAG, hallucination, conversation quality, and more. LLM-as-Judge\ \ Define custom grading criteria and run them with any LLM — GPT-4o, Gemini, Claude, Ollama, or any LiteLLM-supported model. Streaming\ \ Check LLM output token-by-token as it streams. Detect toxic content, PII, or quality drops mid-generation and stop early. Feedback Loops\ \ Submit corrections to scoring results, calibrate thresholds over time, and store feedback in ChromaDB for continuous improvement. Datasets\ \ Create, populate, and manage datasets for evaluation. Upload CSV/JSON files, import from HuggingFace, add LLM-generated columns, and run evaluations at scale. Tracing\ \ Set up OpenTelemetry tracing across Python, TypeScript, Java, and C#. Auto-instrument 45+ frameworks or create custom spans with FITracer. Protect\ \ Guard AI inputs and outputs in real-time. Check for content moderation, bias, security threats, and data privacy violations. Knowledge Base\ \ Upload documents to build knowledge bases for RAG evaluation and context injection. Create, update, and manage files. Annotation Queues\ \ Reference for the AnnotationQueue class in the Future AGI Python SDK. Prompt Optimization\ \ Automatically improve your prompts with 6 SOTA algorithms. Random Search, Bayesian, ProTeGi, Meta-Prompt, PromptWizard, and GEPA. Simulation Testing\ \ Test voice AI agents at scale with simulated customer personas. Run conversations, capture audio, and score performance. Introduction\ \ Complete REST API reference for the Future AGI platform. Health Check\ \ Returns 200 status when server is up and running. No authentication required. Get Evals List\ \ Retrieves a list of evaluations for a given dataset, with options for filtering and ordering. Create Eval Group\ \ Creates a new evaluation group within the user's workspace. List Eval Groups\ \ Retrieves a paginated list of evaluation groups for the user's workspace, including sample groups. Retrieve Eval Group\ \ Retrieves detailed information about a specific evaluation group, including its members. Update Eval Group\ \ Updates an entire evaluation group's details. Delete Eval Group\ \ Soft deletes an evaluation group and removes all its associated evaluation templates. Apply Eval Group\ \ Applies an evaluation group to a set of data, creating user evaluation metrics. Edit Eval List\ \ Adds or removes evaluation templates from an evaluation group. Get Eval Log Details\ \ Retrieves detailed logs for a specific evaluation template, with support for advanced filtering, sorting, and pagination. This endpoint uses a GET req... Create Scenario\ \ Creates a new scenario from a dataset, a script, or a generated/provided graph. The creation is processed in the background. Edit Scenario\ \ Updates the properties of a specific scenario, such as its name, description, associated graph, or the simulator agent's prompt. Add Empty Rows\ \ Adds a specified number of empty rows to an existing scenario. This is useful for populating a scenario with placeholders for future data entry. Add Rows with AI\ \ Initiates an asynchronous task to generate and add a specified number of new rows to a scenario's dataset using AI. A description can be provided to g... Create Agent Definition\ \ Create a new agent definition and its first version. Create Agent Version\ \ Create a new version of an existing agent definition by providing updated agent properties and a commit message. Create Run Test\ \ Creates and configures a new test run, associating it with scenarios, an agent definition, and detailed evaluation configurations. Execute Run Test\ \ Triggers the execution of a specified test run. The execution can be customized to include or exclude specific scenarios. Create Dataset\ \ Create a new dataset with rows and columns in your organization. Upload Dataset from File\ \ Create a new dataset by uploading a local file. Create Score\ \ Create a single annotation score on a source. Bulk Create Scores\ \ Create multiple scores on a single source in one request. Get Scores for Source\ \ Retrieve all scores for a specific source. List Scores\ \ List scores with optional filters. Delete Score\ \ Soft-delete a score. Only the creator or org admin can delete. Create Label\ \ Create a new annotation label. List Labels\ \ List annotation labels with optional filters. Get Label\ \ Retrieve a specific annotation label by ID. Update Label\ \ Update an existing annotation label. Delete Label\ \ Soft-delete an annotation label. Restore Label\ \ Restore a previously deleted annotation label. Create Queue\ \ Create a new annotation queue with assignment strategy and configuration. List Queues\ \ List annotation queues with optional filtering and pagination. Get Queue\ \ Retrieve details of a specific annotation queue. Update Queue\ \ Update an existing annotation queue's configuration. Delete Queue\ \ Soft-delete an annotation queue. Update Status\ \ Transition an annotation queue to a new status. Get Progress\ \ Retrieve progress statistics for an annotation queue. Get Analytics\ \ Retrieve detailed analytics for an annotation queue. Get Agreement\ \ Retrieve inter-annotator agreement metrics for a queue. Export\ \ Export annotation queue items and their annotations as JSON or CSV. Export to Dataset\ \ Export completed annotations from a queue into a FutureAGI dataset. Add Label to Queue\ \ Attach an annotation label to a queue. Remove Label\ \ Detach an annotation label from a queue. Get or Create Default\ \ Get the default annotation queue for a project, dataset, or agent, creating one if it doesn't exist. Find Queues for Source\ \ Find annotation queues that contain a specific source item. List Items\ \ List items in an annotation queue with optional filtering and pagination. Add Items\ \ Add source items to an annotation queue in bulk. Bulk Remove Items\ \ Remove multiple items from an annotation queue at once. Get Annotate Detail\ \ Retrieve a queue item with full source data for the annotation UI. Get Next Item\ \ Retrieve the next available item for the current user to annotate. Submit Annotations\ \ Submit annotations and notes for a queue item. Complete Item\ \ Mark a queue item as completed and optionally receive the next item. Skip Item\ \ Skip a queue item, marking it as skipped by the current user. Get Item Annotations\ \ Retrieve all annotations submitted for a specific queue item. Assign Items\ \ Assign queue items to a specific annotator. Release Item\ \ Release a reserved queue item so it can be assigned to another annotator. Bulk Annotate Spans\ \ Submit annotations and notes for multiple observation spans in a single request.
No results found
↑ ↓ Navigate↵ Open
Table of Contents
Why AI Explainability Is Now a Legal and Ethical Requirement Under GDPR and the EU AI Act
Is it really safe to let AI make decisions? We need to know how an AI came to its decision in order to trust its suggestion. After all, we’re people, and we have a right to know what factors went into the AI’s choice.
Explainable AI isn’t just an a fashion statement; in fields like finance, healthcare, and self-driving cars, we can’t afford a model that gives us an answer without showing its work. Mistakes can cost lives or millions of dollars.
Right now, too many AI systems act like black boxes: they deliver a verdict but keep their reasoning locked away, making it impossible to verify or challenge their output when consequences are high.
That’s why regulators are stepping up. Under GDPR, any personal data used by AI must be handled lawfully, fairly, and transparently-which means you have to explain to users how their data contributes to the outcome.
And the EU’s new AI Act goes even further, sorting AI applications into risk tiers so that “high-risk” systems-think medical diagnosis or credit scoring-follow strict transparency mandates, like documenting decision paths and opening up to third-party audits.
In short, to rely on AI, we need to pull back the curtain. That means building explainability into every model, following GDPR’s “fairness and transparency” mantra, and meeting the AI Act’s accountability bar for high-stake systems-only then can AI truly earn our trust.
What Is AI Explainability? Interpretability vs Explainability Explained
It is very important to know how models make decisions in AI. Explainability and interpretability are two critical concepts in this context.
Interpretability:
- It tries to figure out how an AI model makes its predictions.
- Involves understanding the internal mechanics of the model.
- Is designed to ensure that the model’s operations are clear to humans.
Explainability:
- Tries to figure out why an AI model generated a certain prediction.
- Requires the provision of rationale for the model’s outputs.
- Is designed to ensure that the model’s decisions are understandable to humans.
It is essential to understand these distinctions to be able to create AI systems that are both transparent and reliable.
Technical Challenges in Building Transparent AI: Why Black-Box Models Are a Problem
Developing transparent AI models comes with challenges. Complex models, such as deep neural networks, frequently function as “black boxes,” rendering it challenging to comprehend their decision-making processes. This lack of transparency can result in bias, a lack of accountability, and a decrease in user trust. To solve these problems, we need to create tools and methods that make things more clear so that everyone can understand, trust, and handle AI systems well.
Understanding Black-Box AI Models: Deep Learning, Ensemble Methods, and Built-In Explainability
It becomes increasingly difficult to understand the decision-making processes of AI systems as they become more complex. This section looks into the complex structure of modern AI architectures and the significance of increasing their transparency.
Deep Learning Models
Powerful tools in AI include deep learning models like graph-based models, Transformers, CNNs, and RNNs. Unfortunately, their complexity makes them hard to understand.
- CNNs: Designed mostly for image processing, CNNs include several layers that extract data-based characteristics. It is difficult to understand the extent to which every layer contributes to the ultimate decision.
- RNNs: Designed for sequential data-text or time series-RNNs preserve information across sequences. It is more difficult to determine decision paths due to their recursive nature.
- Transformers: These models are capable of capturing long-range dependencies and managing large datasets. They are effective, but their attention mechanisms introduce layers of complexity that mask their inner workings.
- Graph-Based Models: Complex relationships between data elements are managed by graph-based models, which operate on graph structures. The complex relationships they establish complicate the understanding of their decision-making processes.
The complex, multi-layered structures of these models make them “black boxes,” which makes them very hard to understand.
Non-Traditional Models
Bagging and boosting techniques among ensemble methods combine several models to raise performance. Despite their effectiveness, they pose particular explainability difficulties.
- Bagging: This method generates several separate models averaging their predictions. The diversity among models impacts knowledge of the general decision-making process.
- Boosting: AdaBoost and Gradient Boosting repeatedly create models that fix errors. The layered adjustments complicate the traceability of individual predictions formation.
The complexity of these ensemble techniques often makes it difficult to understand how they get at certain conclusions.
Why Built-In Explainability Matters
AI models must be balanced in performance with openness. Although complex algorithms may show great accuracy, their opacity may reduce responsibility and confidence.
- Trade-offs: While highly complex models may provide superior performance, they are not transparent. It is possible that some accuracy may be sacrificed in exchange for the clarity that simpler, inherently interpretable models provide.
- Inherently Interpretable Models vs. Post-Hoc Methods: Inherently interpretable models are intended to be transparent from the outset, providing a clear understanding of their decision-making processes. On the other hand, post-hoc methods try to explain choices that have already been made, which may not be as accurate.
Focusing on built-in explainability keeps AI systems productive and trustworthy, especially in decision-making applications.
Advanced Techniques for AI Explainability: Post-Hoc Methods, Feature Attribution, and Intrinsic Design
Post-Hoc Explanation Methods
Post-hoc approaches are designed to explain on AI model predictions after they have already been generated, offering valuable insights without modifying the original models.
Model-Agnostic Techniques
LIME (Local Interpretable Model-Agnostic Explanations)
LIME provides an explanation for individual predictions by locally approximating the complex model with an interpretable one. It modifies the input data, tracks how predictions vary, and applies a basic model to these variances in order to comprehend the decision-making process. However, identifying a relevant disturbance zone in high-dimensional spaces is difficult, which may compromise explanation reliability.
SHAP (SHapley Additive Explanations)
SHAP uses ideas from cooperative game theory, especially Shapley values, to give each feature a number that tells how important it is for a certain prediction. This method ensures a fair allocation of the “payout” (prediction) among characteristics, therefore offering consistently locally correct explanations. Deep learning and ensemble approaches are among the several advanced model types SHAP has been modified to handle, hence improving its relevance across different AI systems.
Counterfactual Explanations
Counterfactual models find the least necessary modifications in the input data that impact the prediction of the model. It helps users understand decision limits and how sensitive models are to certain traits by creating these different situations. When judging the quality of a counterfactual, things like its closeness (how similar it is to the real situation), validity (whether it leads to the desired outcome), and plausibility (how likely it is that the unreal event will happen) must be taken into account.
Feature Attribution Methods
Integrated Gradients & Layer-Wise Relevance Propagation
Integrated Gradients connect the model’s output to its inputs by combining the model’s output’s gradients along a path from a baseline to the real input. This approach solves problems like gradient saturation, in which very zero values could cause conventional gradients to miss feature relevance. The prediction is sent backwards through the network layers by Layer-Wise Relevance Propagation (LRP). Each neuron is given a relevance score, which is then added up to find out what input features were contributed.
Permutation Importance and Sensitivity Analysis
Permutation Importance figures out how important each feature is by checking how much the model fails when the values of each feature are mixed up randomly. This approach evaluates the model’s individual feature dependence. Sensitivity Analysis shows how input features impact model output, revealing model robustness and stability. Particularly with big datasets and complex models, both approaches need careful attention of statistical robustness and computing efficiency.
Advanced methods like these make it easier to understand and trust AI models, which supports their secure and helpful use in many areas.
Intrinsic Explainability Techniques
Intrinsic explainability targets the development of AI models that are naturally comprehensible, which reduces the need for external explanation methods.
Interpretable Model Design
Some models have transparent nature because of their design:
- Decision Trees: These models depict decisions as a tree structure, with each node representing a feature and each branch representing a decision rule. This style makes it easy for users to see how information goes from input to output, which helps them make decisions.
- Rule-Based Systems: These systems make decisions based on clear “if-then” rules that have already been determined.
- Attention Mechanisms: In models like as transformers, attention mechanisms underline which areas of the input data the model concentrates on during processing. This mechanism gives you information about how the model makes decisions.
Combining these interpretable models with complex ones can generate hybrid architectures balancing performance with clarity.
Mathematical Optimization for Interpretability
Mathematical methods improve model transparency:
- Regularization Methods: Lasso (Least Absolute Shrinkage and Selection Operator) is a technique that adds a penalty to the model for the use of too many features. This encourages simplicity and makes the model’s decisions simpler to interpret.
- Multi-Objective Optimization Frameworks: These frameworks are designed to achieve a balance between the accuracy of the model and its interpretability by optimizing both objectives simultaneously.
Using these techniques together, AI experts can create models that work well and are easy to understand, which builds trust and understanding in AI systems.
Figure 1: Intrinsic Explainability Techniques in AI
Modern Explainability in Large Language Models
Traditional post-hoc explanation techniques, such sensitivity analyses and feature attribution, have proven to be quite successful for classical machine learning models. But as LLMs develop, their transparent, emergent reasoning capacity requires new explainability models. Modern methods now take advantage of the natural ability of the model to produce its own explanations together with solutions. This method of self-generated explanation is meant to make things clearer by asking the model to describe a “chain of thought” (CoT) that shows the steps it took to arrive at its end result.
LLM-Generated Explanations
Asking the LLM to provide an explanation along with its response is a promising strategy. For example, the model can be driven instead of providing a simple numerical answer for a mathematical problem:
“Solve the problem and explain step by step how you arrived at the solution.”
This approach lets users examine the intermediate thinking phases and evaluate if the justification makes sense and fits the last result. However, research have shown that these produced explanations don’t always accurately capture the underlying workings of the model. In many cases, the explanation may be logical, but it may not accurately represent the true factors that influence the answer or may even deviate from the actual decision pathway. This phenomenon is frequently referred to as “ unfaithful explanations.” So, the answer might be right, but the reason that goes with it might not be a clear look into the model’s “thoughts,” but rather a justification for what happened.
Real-World Examples: Large language models (LLMs) have become more useful than only code generators thanks to recent developments. Debugging assistance is now frequently provided by modern platforms through the provision of natural-language explanations for code behavior. Recent research has shown that LLMs that have been trained on code, such as those in GitHub Copilot, are capable of producing step-by-step narratives that explain the reasoning behind the execution of specific functions and the processing of data. Minor logical errors could go unnoticed when the text doesn’t include specific specifics of the debugging procedure, even though these explanations are usually easy to understand. This research underscores the potential and challenges associated with employing automated explanations in debugging tasks.
Chain-of-Thought Prompting
Chain-of-thought prompting has become the most common strategy for encouraging LLMs to break down complex tasks into sequential reasoning stages. Clearly instructions, such “Let’s think step by step,” help the model to express intermediate processes before offering a final response. This approach offers a structured reasoning trail in addition to improving performance on benchmarks of arithmetic, symbolic, and common sense thinking.
Even with these benefits, there are still some problems. To begin, CoT outputs are supposed to show how the thinking works, but there is more and more proof that they may not be completely accurate. Sometimes the chain of thoughts generated acts as a “story” meant to justify the response by hiding the actual underlying calculations. As so, the explanation could occasionally oppose the response, creating a misleading impression of openness. For disciplines like AI safety and regulatory compliance, where knowledge of the decision-making process is critical, this restriction presents serious challenges.
Additionally, the repetitive nature of chain-of-thought reasoning-in which models can change earlier steps or come up with multiple possible paths-brings up questions about which chain, if any, correctly shows how the model works on the inside. Researchers are looking into ways to make things more reliable, like self-consistency decoding, which combines different CoT outputs. Nevertheless, these methods continue to come across challenges in guaranteeing that the visible chain-of-thought accurately reflects the concealed inference process.
Real-World Examples: Chain-of-thought prompting has been identified as an effective method for assisting students in the resolution of complex mathematical problems in educational environments. One of the most recent studies showed that the solution process can be rendered more transparent by encouraging LLMs to deconstruct multi-step math problems into intermediate reasoning steps. However, classroom experiments have shown that the generated explanations are logical but the final numerical answers occasionally do not align precisely with them. This discrepancy implies that a portion of the model’s explanation may function as a post-hoc rationalization rather than a precise representation of its internal reasoning.
Interactive and Adaptive Explanation Interfaces
Modern LLMs have been specifically targeted by interactive explanation interfaces that have emerged as a result of recent advancements in explainable AI. The development of interactive explanation interfaces for complex language models, particularly in the domain of recommender systems, has been a result of recent advancements in explainable AI. Recent research conducted by Tintarev and colleagues showed that adaptable interfaces can enable end-users to not only observe but also interact with AI-generated explanations. For example, a 2023 co-design study for multi-stakeholder job recommender system explanations demonstrated that interactive features-such as the capacity to query, refine, and dispute explanations-can be effectively customized to meet the diverse requirements of candidates, recruiters, and companies. In the same manner, research on group recommender system explanations has highlighted the importance of creating visual and textual interfaces that enable users to actively manage the explanation process, which promotes trust and transparency.
Interactive components help systems to provide various advantages:
- Improved Trust: Users have the ability to explore the reasoning steps to understand the process by which recommendations or decisions were made.
- Feedback Loops: Adaptive interfaces enable human-in-the-loop feedback, which enables the iterative enhancement of both the content of the explanation and the underlying model.
- Customization based on Context: The system can change answers based on the user’s knowledge and the situation, which makes the result easier to understand and use.
- Complex decision-making systems in important fields like healthcare and finance require this interactive approach since understanding and confirming an AI’s logic could impact user trust and safety.
Real-World Examples: Next-generation recommender systems will likely have interactive explaining displays as standard. It focuses on the empowerment of users by offering multi-level, adjustable explanations. For instance, a recent study provided a framework that provides both retrospective and prospective explanations through counterfactual reasoning, enabling users to interactively adjust parameters to gain a deeper awareness of the reasons why items are recommended. Another study developed an explainable scientific literature recommender system that allows users to select from a variety of levels of explanation detail (basic, intermediate, and advanced). In addition to enhancing transparency and trust, these studies also increase user satisfaction by providing users with more control over the recommendation process.
Although self-generated chain-of-thought prompting and interactive explanation interfaces have made major advances in demystifying LLM reasoning, the importance of rigorous, quantifiable metrics for AI explainability is underscored by the persistent challenges in ensuring that these explanations accurately reflect internal decision-making. We will explore these metrics in the following section.
Metrics for Evaluating AI Explainability
Evaluating AI explainability ensures models’ transparency and dependability by use of both quantitative and qualitative evaluations.
Quantitative Metrics
These metrics offer statistical evaluations of the effectiveness of an explanation.
- Fidelity and Consistency: This measures the actual behavior of the model reflected in an explanation. High accuracy means that the description closely matches how the model makes decisions.
- Stability and Sensitivity: This evaluates the robustness of explanations to variations in the supplied data. Stable explanations are consistent across similar inputs, while sensitivity analysis evaluates the robustness of explanations when inputs are disturbed.
- Complexity Metrics: This assesses the equilibrium between the comprehensiveness and simplicity of explanations. Simpler explanations are easier to grasp, but they have to match the model’s behavior.
Qualitative Assessments
Qualitative assessments emphasize the clarity and effectiveness of explanations by involving human evaluators. Human-in-the-loop reviews involve users to find out if the answers help them understand, build trust, and make decisions. These methods ensure that explanations are technically solid, insightful, and user-friendly.
Combining these metrics helps professionals to evaluate and enhance the explainability of AI models which ensures their accuracy and efficiency.
AI Explainability Tools and Frameworks in 2026
AI explainability has been improved in 2025 by the development of several advanced tools and frameworks that assist developers and researchers in understanding complex models.
Open-Source Libraries
Captum (for PyTorch)
Captum is a comprehensive library that is specifically designed for PyTorch models. It provides a variety of attribution methods to assist in the interpretation of model predictions.
- Integrated Gradients: You can find out how important a trait is by figuring out the integral of slopes with respect to inputs along a path from a baseline to the actual input.
- DeepLIFT (Deep Learning Important Features): DeepLIFT provides a speedier alternative to Integrated Gradients by comparing the activation of each neuron to its reference activation and assigning contribution scores.
- Custom Attribution Methods: Captum allows the creation of custom attribution techniques so that users can customize explanations to certain model architectures and needs.
Alibi
Alibi is a Python library that specializes in the investigation and interpretation of machine learning models, offering high-quality implementations of a variety of explanation methods.
- Use Cases: Alibi provides approaches suitable for tabular, text, and picture data that fit both classification and regression models.
- API Examples: The library makes it simple to include into current machine learning pipelines as it offers a consistent API for several explainers.
- Integration: Alibi can be easily included into several deployment systems, which allows the production environment deployment of explanations.
Platforms
To improve AI explainability, tools have been created alongside open-source libraries.
Google’s What-If Tool
What-If Tool, developed by Google, enables users to interactively investigate machine learning models.
- Interactive Model Interrogation: Users can alter input data points and observe the impact of these changes on model predictions, which is helpful in comprehending model behavior.
- Scenario Analysis: The tool facilitates the examination of alternative scenarios, which assists in the identification of potential biases and the evaluation of the robustness of the model.
IBM’s AI Explainability 360
IBM’s AI Explainability 360 is a comprehensive framework that provides a variety of algorithms to elucidate machine learning models.
- Comparative Analysis: The toolkit offers a variety of explainability techniques, therefore enabling users to evaluate many ways and choose the best suitable one for their particular requirements.
- Open-Source Alternatives: AI Explainability 360 is a proprietary product that supports open-source libraries such as Captum and Alibi by providing a user-friendly interface and additional methods.
These tools and frameworks are very important for making AI models less mysterious so that people in many different fields can trust them more.
Challenges in AI Explainability: Scalability, Oversimplification, and Explanation Instability
Creating AI systems that can be understood by humans is not an easy task. Scalability is a challenging task in deep neural networks because of their complex architectures. Simplifying these models to improve understanding can result in a decrease in performance, highlighting the fundamental trade-off between transparency and accuracy.
Limitations are also present in the existing explainability methods. The potential for misunderstanding exists when explanations that fail to convey the model’s intricacies are the result of oversimplification. Furthermore, the stability and reliability of the explanations they offer can be impacted by the sensitivity of many current methods to minor changes in input data.
Why Transparent AI Is Essential for Trust, Compliance, and Widespread Adoption Beyond 2026
We’ve talked about advanced methods and tools that make AI easier to understand, like model-agnostic methods like LIME and SHAP and intrinsic approaches like interpretable model design. To evaluate the effectiveness of these justifications, we also looked at evaluative criteria like complexity, stability, and authenticity.
Transparent AI is important for long-term progress. It makes sure that AI systems are reliable, responsible, and in line with human ideals, which is very important for their widespread use in many areas.
Looking ahead, beyond 2025 we expect major developments in explainability methods. Explainable Reinforcement Learning is an emerging area that is garnering attention due to its objective of enhancing the transparency of the decision-making process in reinforcement learning. Likewise, work is under progress to improve Graph Neural Network interpretability used to replicate complex relational data. More clear and easily available AI systems resulting from these advances could help to build more confidence and wider acceptance in society.
Future AGI goes beyond basic explainers by embedding real-time evaluation tools like Observe for continuous monitoring and Protect for automated risk mitigation, so you can keep an eye on toxicity, bias, and safety as your models run in production.
Frequently Asked Questions About AI Explainability Tools and Techniques
What is the difference between interpretability and explainability in AI?
Interpretability is how easy it is for someone to look inside a model. For example, a decision tree lets you see exactly how inputs are linked to outputs. Explainability, on the other hand, is more about giving a clear, easy-to-understand reason for why a certain prediction was made, even if the model’s inner workings are still hard to understand.
What are the most common post-hoc explanation methods for AI models?
With LIME, you take a black‐box model and approximate its behavior around a single prediction by training a simpler, local model-kind of like zooming in on one small area to see what’s happening there. SHAP uses ideas from game theory (Shapley values) to assign every feature a fair “importance score” for a particular prediction, so you can see which inputs played the biggest role.
What are intrinsic interpretability techniques in machine learning?
From the beginning, this model is transparent. For instance, decision trees show you clear paths to follow, rule-based systems use simple “if-then” logic, and methods like Lasso regression penalize unnecessary features so the final model stays simple and easy to understand.
What modern methods exist for explaining large language models?
When you use Chain-of-Thought prompting with an LLM, you ask the model to walk you through its reasoning step by step-”Let’s think this through together,” essentially-before it gives you a final answer. Just keep in mind that sometimes these chains sound more like post-hoc justifications (what the model wants you to see) rather than the exact internal calculations happening under the hood.
Related Articles
\ \ Guides](/content/blog/ai-chatbot-guide-2025/index.html)
Step-by-Step Guide on Building Generative AI Chatbot 2025
Learn how to build a generative AI chatbot in 2026. Covers LLM selection, RAG pipelines, evaluation metrics, real-time monitoring & safety guardrails.
Rishav Hada·Jul 24, 2025
5 min
\ \ Guides](/content/blog/top-5-llm-evaluation-tools-2025/index.html)
Top 5 LLM Evaluation Tools of 2025
Compare top 5 LLM evaluation tools in 2026. Covers Future AGI, Galileo, Arize, MLflow, and Patronus AI across capabilities, scalability, and use cases.
Rishav Hada·Apr 30, 2025
5 min
\ \ Guides](/content/blog/understanding-langchain-callback-how-to-use-it-effectively/index.html)
Understanding Langchain Callback: How to Use It Effectively
Learn how Langchain callbacks work in 2026. Covers core callback events including on_chain_start and on_tool_end, built-in vs custom callback handlers.
Ashhar Aziz·Mar 7, 2025
5 min
Stay updated on AI observability
Get weekly insights on building reliable AI systems. No spam.
Subscribe
You're subscribed!
AI AssistantBeta
FutureAGI AI Assistant
Ask me anything about the FutureAGI platform — I can search across all docs instantly.
What is FutureAGI?What can FutureAGI do?How do I run my first evaluation?How do I set up tracing?How do I detect hallucinations?
Built by FAGI with ❤️
Explain "AI Explainability in 2026: …"