Sandbox
@growthxai/output

TypeScript framework for AI workflows and agents

Output is a TypeScript framework that gives you workflows, steps, prompts, evaluators, tracing, and credentials in one place. It is built so Claude Code can work inside the same repo, with each workflow kept in one folder alongside its code, prompts, tests, and evals. The stack uses Temporal for durable execution and supports multiple LLM providers through one API.

435 stars12 forksJavaScriptUpdated 6d ago
Who it's for

Builders who want Claude Code or Codex to build, test, and iterate on AI workflows in their codebase.

What it delivers

You can ship agent-led workflows with retries, tracing, evals, and prompt management already wired in.

What it does

Workflow and step primitives

Defines deterministic workflows and side-effecting steps as separate building blocks.

Prompt files

Stores prompts in `.prompt` files with YAML frontmatter and Liquid templating, versioned with code.

Evaluators

Runs LLM-as-judge, offline, and inline evaluators with confidence scores and reasoning.

Tracing and cost tracking

Captures LLM calls, HTTP requests, token counts, costs, latency, and full prompt and response pairs.

Multi-provider LLM support

Works with Anthropic, OpenAI, Azure, Vertex AI, and Bedrock through one API.

Temporal-based execution

Uses Temporal for retries, replay, child workflows, parallel execution, and workflow history.

Encrypted credentials

Stores API keys and other secrets with AES-256-GCM and CLI-managed environment scoping.

CLI and local dev stack

Provides project init, local dev, workflow run, and debug commands.

How to get it

  1. 1Scaffold a project and add your API key to .env (ANTHROPIC_API_KEY=sk-ant-...)
    npx @outputai/cli init
    cd <project-name>
  2. 2Start the full development environment — Temporal server, API server, a worker with hot…
    npx output dev
  3. 3Run your first workflow and inspect the execution
    npx output workflow run blog_evaluator paulgraham_hwh
  4. 4Run
    npx output workflow debug <workflow-id>

README

Output

GitHub stars npm downloads License: Apache-2.0 TypeScript Build Status

The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code — describe what you want, Claude builds it, with all the best practices already in place.

One framework. Prompts, evals, tracing, cost tracking, orchestration, credentials. No SaaS fragmentation. No vendor lock-in. Everything in your codebase, everything your AI coding agent can reach.

Output.ai Demo
Watch a complete example of using Output to build a newsletter pipeline

Why Output

Every piece of the AI stack is becoming a separate subscription. Prompts in one tool. Traces in another. Evals in a third. Cost tracking across five dashboards. None of them talk to each other. Half of them will get acquired or shut down before your product ships.

Output brings everything together. One TypeScript framework, extracted from thousands of production AI workflows. Best practices baked in so beginners ship professional code from day one, and experienced AI engineers stop rebuilding the same infrastructure.

Build AI using AI

Output is the first framework designed for AI coding agents. The entire codebase is structured so Claude Code can scaffold, plan, generate, test, and iterate on your workflows. Every workflow is a folder — code, prompts, tests, evals, traces, all together. Your agent reads one folder and has full context.

Own your prompts

.prompt files with YAML frontmatter and Liquid templating. Version-controlled, reviewable in PRs, deployed with your code. Switch providers by changing one line. No subscription needed to manage your own prompts.

See everything that happens

Every LLM call, HTTP request, and step traced automatically. Token counts, costs, latency, full prompt/response pairs. JSON in logs/runs/. Zero config. Claude Code analyzes your traces and fixes issues — because the data is in your file system.

Test AI like software

LLM-as-judge evaluators with confidence scores. Inline evaluators for production retry loops. Offline evaluators for dataset testing. Deterministic assertions and subjective quality judges.

Use any model

Anthropic, OpenAI, Azure, Vertex AI, Bedrock. One API. Structured outputs, streaming, tool calling — all work the same regardless of provider.

Scale without worrying

Temporal under the hood. Automatic retries with exponential backoff. Workflow history. Replay on failure. Child workflows. Parallel execution with concurrency control. You don't think about Temporal until you need it — then it's already there.

Keep secrets secret

AI apps need a lot of API keys. Sharing .env files is risky, and coding agents shouldn't see your secrets. Output encrypts credentials with AES-256-GCM, scoped per environment and workflow, managed through the CLI. No external vault subscription needed.

Quick Start

Requirements:

Scaffold a project and add your API key to .env (ANTHROPIC_API_KEY=sk-ant-...):

npx @outputai/cli init
cd <project-name>

Start the full development environment — Temporal server, API server, a worker with hot reload, and the Temporal UI at http://localhost:8080:

npx output dev

Run your first workflow and inspect the execution:

npx output workflow run blog_evaluator paulgraham_hwh
npx output workflow debug <workflow-id>

For the full getting started guide, see the documentation.

Core Concepts

Workflows

Orchestration layer — deterministic coordination logic, no I/O.

// src/workflows/research/workflow.ts
workflow({
  name: 'research',
  fn: async (input) => {
    const data = await gatherSources(input);
    const analysis = await analyzeContent(data);
    const quality = await checkQuality(analysis);
    return quality.passed ? analysis : await reviseContent(analysis, quality);
  }
});

Steps

Where I/O happens — API calls, LLM requests, database queries. Each step runs once and its result is cached for replay.

// src/workflows/research/steps.ts
step({
  name: 'gatherSources',
  fn: async (input) => {
    const results = await searchApi(input.topic);
    return { sources: results };
  }
});

Prompts

.prompt files with YAML configuration and Liquid templating.

---
provider: anthropic
model: claude-sonnet-4-20250514
temperature: 0
---

<system>You are a research analyst.</system>
<user>Analyze the following sources about {{ topic }}: {{ sources }}</user>

Evaluators

LLM-as-judge evaluation with confidence scores and reasoning.

// src/workflows/research/evaluators.ts
evaluator({
  name: 'checkQuality',
  fn: async (content) => {
    const { output } = await generateText({
      prompt: 'evaluate_quality',
      variables: { content },
      output: Output.object({
        schema: z.object({
          isQuality: z.boolean(),
          confidence: z.number().describe('0-100'),
          reasoning: z.string()
        })
      })
    });

    return new EvaluationBooleanResult({
      value: output.isQuality,
      confidence: output.confidence,
      reasoning: output.reasoning
    });
  }
});

SDK Packages

PackageDescription
@outputai/coreWorkflow, step, and evaluator primitives
@outputai/llmMulti-provider LLM with prompt management
@outputai/httpHTTP client with tracing
@outputai/cliCLI for project init, dev environment, and workflow management

Example Workflows

Production-ready workflows you can run locally, learn from, and fork — all from the output-examples gallery:

WorkflowDescriptionAPIs
blog_evaluatorEvaluate blog post signal-to-noise qualityJina Reader
call_scorerScore sales call transcripts against MEDDIC, BANT, or SPINLLM only
changelog_generatorGenerate categorized changelogs from GitHub commits and PRsGitHub
dependency_auditAudit npm dependencies for vulnerabilities, licenses, and abandonmentGitHub, OSV, npm
recipe_extractorExtract structured recipes from blog URLsJina Reader
url_summarizerSummarize any webpage into TLDR, key points, and FAQJina Reader
youtube_summarizerSummarize YouTube videos with key moments and takeawaysYouTube
ai_hn_digestPersonalized Hacker News digest published to Beehiiv newsletterHN, Jina Reader, Beehiiv
sales_call_processorProcess sales call transcripts into notes + parallel recipe analysesLLM only

Browse the full gallery at output.ai/gallery.

Projects using Output

ProjectDescription
CheckThatCheckThat is an AEO platform built on Output's durable, deterministic LLM workflows — tracking how B2B brands show up across ChatGPT, Claude, Perplexity, and Google AI, covering 2.6M+ AI responses spanning 5,875+ brands.

Configuration

For production configuration and advanced settings (LLM providers, Temporal Cloud, tracing, and more), see the operations docs.

Contributing

See CONTRIBUTING.md.

License

Apache 2.0 — see LICENSE file.

Acknowledgments

Built with Temporal, Vercel AI SDK, Zod, LiquidJS.

Files in the repo

Repository payload36 top-level entries
  • .changeset
  • .claude
  • .claude-plugin
  • .github
  • .husky
  • api
  • assets
  • coding_assistants
  • docs
  • ops
  • sdk
  • test_workflows
  • .dockerignore
  • .gitignore
  • .npmrc
  • .tool-versions
  • CLAUDE.md
  • CODE_OF_CONDUCT.md
  • context7.json
  • CONTRIBUTING.md
  • docker-compose.dev.yml
  • docker-compose.prod.yml
  • docker-compose.temporal.yml
  • eslint.config.js
  • LICENSE
  • package.json
  • pnpm-lock.yaml
  • pnpm-workspace.yaml
  • README.md
  • RELEASING.md
  • render.yaml
  • run.sh
  • SECURITY.md
  • typedoc.json
  • vitest.config.js
  • vitest.integration.config.js

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More frameworks & sdks

HKUDS/nanobotFrameworks & SDKs

Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps

48k
microsoft/
SkillOpt
microsoft/SkillOptFrameworks & SDKs

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

17k
omnigent-ai/omnigentFrameworks & SDKs

Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.

9.8k
kyegomez/
OpenMythos
kyegomez/OpenMythosFrameworks & SDKs

A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.

15k
D4Vinci/ScraplingFrameworks & SDKs

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

80k