← Knowledge

Letta Office Hours: Hosted MCP, Workflows, and Agent Orchestration

Hosted MCP agents in Claude Code, parallel workflows, model and memory updates, and the limits of agent orchestration.

Watch on YouTube ↗

The September 24, 2026 Letta Office Hours shows two ways to extend a persistent agent's reach. A hosted Model Context Protocol (MCP) server lets compatible clients, including Claude Code, call Letta agents. The Workflow tool lets an agent divide a bounded job among parallel subagents and combine their results. The recording also covers model choices, community mods, memory migration, and the division of labor between humans and agent-operated software.

This episode guide records what was shown or discussed that day. The live Q&A often distinguishes a working demo from a proposed feature. That distinction matters especially for MemFS v2 and automatic tool approvals. Use the recording and current product documentation before configuring a production integration.

Selected chapters

Time Topic
00:00 Hosted MCP server and Claude Code demo
04:26 Workflow tool and parallel subagents
08:07 Model roundup
12:03 Grok authentication on Letta Cloud
14:15 Community Sprite mod
18:02 Experimental desktop and web voice transcription
27:59 MemFS v2 and migration status
29:22 Workflows compared with the Agent SDK
01:01:32 Unified toolset experiment
01:05:40 Channels compared with custom SDK relays
01:10:17 Human judgment in agent-assisted development
01:40:29 Agents as orchestrators
01:43:47 Structured choices for tool approvals

An agent can be a tool in another client

The hosted MCP server exposes Letta agents to software that accepts an MCP URL and API key. Cameron demonstrates using an agent from Claude Code rather than switching back to the agent's usual interface. The six tools discussed cover listing agents and models, creating an agent, sending a message, and retrieving a run or reply. In the episode, the hosted integration uses an API key; OAuth support is discussed as future work. Cameron expresses uncertainty about whether this path currently reaches local agents, so the demo should not be read as a guarantee of local execution.

MCP provides an interface for invoking the agent. The agent still brings its own memory and conversation, while the calling client supplies a task and receives a result. That makes an existing agent usable inside a coding environment without turning the coding environment into the agent's only home.

The demo connects the MCP server in a client, discovers an existing agent, sends it a message, then retrieves the run and its answer. The client can also create a new agent through the same interface. Listing, messaging, and run retrieval are separate operations rather than a single opaque chat endpoint. This lets a client distinguish an accepted request from a completed reply.

API-key support limits which MCP clients can connect. A client that only supports OAuth cannot use this route yet. The server's reach also depends on where the target agent executes. Cameron says he believes the current path supports cloud agents, while leaving local execution uncertain. The Claude Code demo does not establish broader access.

Workflows split bounded jobs into parallel work

The Workflow tool runs a script that dispatches subagents, then collects their findings. Cameron shows an example that researches note-taking applications in parallel and synthesizes a comparison. He also suggests code review and trip research as candidate tasks. These are demonstrations of a pattern, not verified results for every proposed use case.

In the note-taking demo, the primary agent writes a workflow, assigns different applications to three researchers, and gives their results to a synthesis worker. Parallelism is useful because the researchers can investigate independent candidates at the same time. A code review could use a different split: have workers inspect security, correctness, and performance, then verify and prioritize the findings. The second stage matters. Three confident reviews do not become reliable merely because they agree in a summary.

The distinction from the Agent SDK is lifetime and ownership. A workflow script coordinates a bounded run; an SDK application can retain routing, integrations, and ongoing behavior beyond that run. Cameron also distinguishes these short-lived workflow workers from messaging another persistent agent with its own history. Sending every small question through multiple workers adds overhead and model cost without necessarily improving the answer.

The agent's choice to delegate is itself a design decision. A job with independent research branches and an explicit synthesis step can justify parallel workers. A short factual question usually cannot. Cameron notes that a workflow-writing agent may over-delegate, making a task slower and more expensive than doing it directly. He also says users may have to ask for the Workflow tool explicitly while the interface for displaying its progress is still being developed.

Workflows, applications, and channels have different lifetimes

The Agent SDK supports longer-lived applications that receive events, call agents, and choose where work executes. Cameron sketches an application that receives a repository event and routes a bounded task to a suitable model or computer. That is an architectural example, not a deployment demonstrated in this episode. The SDK can target local or cloud environments and use different models; a Workflow tool invocation is closer to a temporary script for one task.

Built-in channels are convenient when a supported integration, such as Telegram, gets an agent reachable quickly. For a new or highly customized messaging surface, Cameron prefers a small SDK relay: receive an incoming message, call the agent, and carry its reply back. The relay must still be hosted and managed. His recommendation concerns flexibility and lifecycle ownership; it is not an announcement that existing channels are going away.

New surfaces still need maturity checks

The model discussion covers Opus 5.5, GPT-6 Luna and Sol, and Grok 4.7. Cameron recommends Luna for inexpensive, well-scoped subagent tasks and praises Opus for legibility while acknowledging its cost. He likes Grok for code and bounded work, but is less convinced it can replace a model already shaping a personal agent. These are experience reports from this episode, not lasting rankings or price guarantees.

For Grok, he demonstrates connecting an X subscription to Letta Cloud through a device-code OAuth flow in the model/provider settings. Model entitlement depends on the user's subscription and current provider terms. He recommends testing a model change in a fresh conversation before moving an established personal agent, where the model's interaction style may matter as much as task performance. The slide says Grok 4.6; Cameron corrects the newly discussed release to 4.7 during the recording.

The community Sprite mod has expanded from a statusline companion to multiple sprites and optional agent-backed speech. An agent-backed companion has its own model use and memory; the choice of what it may observe deserves separate attention. Cameron also discusses backing up agent memory. Current MemFS uses a Git repository, which makes its files portable, but a memory copy alone should not be mistaken for a complete backup of every part of an agent.

Experimental voice transcription is available behind the desktop/web experiments panel discussed in the episode. After enabling it, a microphone control records audio and inserts the transcription into the chat input for review before sending. Cameron says audio goes to OpenAI for transcription, through Letta Cloud by default or through a user-provided OpenAI key. He has not tested the local-backend path. The live demonstration exposes rough edges, including missing feedback during transcription and a clipped recording; the feature is explicitly preliminary. Voice transcription through Signal and Telegram channels is a separate, earlier capability.

MemFS v2 is discussed as a proposed reorganization of agent memory. The proposed layout would flatten a special system-memory folder so root memory files enter the agent's context, and use indexes to navigate other files. An early remark calls it available, but Cameron later clarifies that general migration is not ready and explicitly advises users not to migrate yet. Code visible in a repository is not a migration instruction. He expects to provide guidance when it is ready. This is separate from moving old memory blocks into the existing Git-backed MemFS format.

Tool naming and automatic approvals are experiments

Letta currently adapts some tool names and behavior to the conventions of a model's native coding harness. In the Q&A, Cameron describes an optional unified toolset with more consistent names across models. His hypothesis is that stronger models may no longer need as much harness-specific translation. He does not present comparative benchmarks or announce the unified set as the default. Testing it with representative tasks would establish more than an intuition about model capability.

The episode closes with a discussion of Jev, a model that chooses among predefined answers rather than generating open-ended prose. A community mod uses that kind of classifier to judge whether a proposed tool call should proceed under the available context and policy. For example, an approval system could evaluate the requested tool, its arguments, and a user's instruction, then return a bounded allow-or-deny judgment. Cameron says he has not tried the mod and describes a native auto-approval classifier as an idea under exploration.

A probability attached to a verdict is not proof that the verdict is safe. The policy, input context, threshold, and failure behavior still need evaluation against the actual workflow. Cameron's practical advice is to define success criteria and benchmark the decision task before relying on a cheap classifier as a permissions layer. A classifier's answer also does not expand the permissions that the application is authorized to grant.

Humans still choose the work

Cameron describes agents doing implementation, review, and investigation while humans evaluate whether the work should exist and whether its effects are acceptable. The point emerges from his account of software development: agents can make a change and review it, but a human still decides which user problem deserves attention, asks whether the review found the right issue, and chooses the product's focus. More generated code cannot settle those questions.

He favors cheaper models for narrow delegated tasks when the task and verification are clear, leaving expensive models for synthesis and direction. A persistent orchestrating agent can keep the larger purpose in view while dispatching work. That division is a working approach, not a benchmark proving one model allocation always wins. Delegation also creates new responsibilities: defining worker scope, checking their outputs, and observing whether the external effect actually happened.

MCP lets a persistent agent enter another tool's workflow. Workflows let an agent divide one job. The SDK can carry work across events and computers. None of those interfaces decides which problem is worth solving or whether a result was delivered correctly. Those judgments remain with the person commissioning and checking the work.

Sources

  1. YouTube episode
  2. Sprite mod
  3. Jev announcement

Suggest a correction ↗

Appearance