When one agent isn't enough

Orchestrating secure parallel agent swarms

Pamela Fox

Python Cloud Advocate · Microsoft · pamelafox.org

Part 1 of 4

Introducing the swarm

One question, many parallel researchers

Introducing the swarm Sandboxes 101 Back to the swarm Wrapping up

One question, several independent research tasks

“My family of four in California is deciding whether to replace our gas car (Subaru Forester 2001) with a hybrid version of the Forester in 2026. Compare the total five-year cost, charging practicality, winter range, reliability, insurance, and available incentives. Include the strongest reasons not to switch, identify assumptions that could change the recommendation, and produce a decision checklist.”

Example research plan

What's the five-year cost of the hybrid vs. keeping the 2001?
Which California incentives apply?
How do winter range and reliability compare?
How much would insurance change?
  • Fan-out makes the answer faster and more thorough.
  • But now several agents need compute, permissions, and cleanup at once.

Demo: the research swarm

The research swarm web app running the Forester upgrade sample topic. The live agent architecture shows Planner complete, a parallel research wave on wave 2 with 4 of 4 researchers complete, the Reviewer complete with a dashed 'gaps, more research' route back to the Planner, and the Report Writer waiting for approval.

Swarm architecture

User question ↓
↓ Final report

Sandbox lifecycle: create sandbox → run research → collect result → delete

Part 2 of 4

Sandboxes 101

Create, configure, secure, and keep state

Introducing the swarm Sandboxes 101 Back to the swarm Wrapping up

What is a sandbox?

Sandbox 1 Sandbox 2 Sandbox 3 Processes Memory Filesystem Host time 123 ● Sub-second startup ✕ Deleted when done Outbound network Model API Telemetry Any other site CPU ≤ 2 cores RAM ≤ 4 GiB

Isolated

Its own processes, memory, and filesystem. It can't see or change the host or other sandboxes.

Fast & disposable

Sub-second startup for one task, then gets deleted when the task is done.

Bounded

Gets only the network endpoints, credentials, and CPU and memory you allow for its task.

Azure Container Apps Sandboxes

Fast, isolated, and stateful compute infrastructure on demand.

Execute securely by default

Sandbox isolation for any untrusted workload

Resume instantly

Preserve state across stop and resume, with enterprise controls

Burst to hyperscale

Sub-second start, zero to thousands, no compute cost when idle

The foundation layer used by:

  • GitHub Copilot cloud sandboxes
  • Microsoft Foundry hosted agents
  • Azure Container Apps Express
  • Azure SRE Agent
  • Microsoft Copilot Studio

Anatomy of a sandbox on Azure

SANDBOX GROUPSANDBOXESsandbox-1RunningM — 1 core / 2 GiB / 20 GiBingress: port 8080sandbox-2RunningL — 2 cores / 4 GiB / 40 GiBsandbox-3StoppedM — 1 core / 2 GiB / 20 GiBingress: port 8080sandbox-4StoppedS — 0.5 cores / 1 GiB / 10 GiBSHARED RESOURCESDisk Imagesroot filesystemsSnapshotsmemory + diskData Volumesblob · data diskSecretsegress credential injectionIdentitymanaged identityInternet / VNetIngress proxyPER-SANDBOX POLICIESoff by defaultopt-in per portEgress proxyPER-SANDBOX POLICIESopen by defaultallow / deny host rulescredential injectioninjected credentialsstay outside the sandboxInternet / VNet

Create a sandbox group


resource group 'Microsoft.App/sandboxGroups@2026-02-01-preview' = {
  name: 'my-sandbox-group'
  location: 'westus2'
}

var dataOwner = 'c24cf47c-5077-412d-a19c-45202126392c'

resource access 'Microsoft.Authorization/roleAssignments@2022-04-01' = {
  name: guid(group.id, principalId, dataOwner)
  scope: group
  properties: {
    roleDefinitionId: subscriptionResourceId(
      'Microsoft.Authorization/roleDefinitions', dataOwner)
    principalId: principalId
  }
}
							

Azure CLI + aca CLI

📎 Quickstart

az group create --name my-rg --location westus2

aca sandboxgroup create -g my-rg \
  --name my-sandbox-group \
  --location westus2 \
  -s "$SUBSCRIPTION_ID" --set-config

aca sandboxgroup role create \
  --group my-sandbox-group \
  --role "Container Apps SandboxGroup Data Owner" \
  --principal-id "$PRINCIPAL_ID"
							

Create a sandbox from the portal

Azure portal Create Sandbox form in simple and quick mode: source set to the ubuntu public disk image, resource tier M (1 core, 2 GiB), and an optional name label. The sandbox's browser terminal running a Python process that holds 1 GiB of memory, with top showing it at about 91% CPU and 1.0 GiB resident. Below, the portal's gauges read CPU 1 core at 100%, memory 2 GiB at 53% (1.2 GiB used), and storage 892 MiB of 19.5 GiB.

📎 portal.azure.com

Create a sandbox programmatically


aca sandbox create --disk ubuntu --label name=demo
aca sandbox exec -l name=demo -c "uname -a"
aca sandbox delete -l name=demo --yes
							

Agent Skill · Copilot CLI

📎 Quickstart

/plugin marketplace add microsoft/azure-container-apps
/plugin install sandboxes@Azure-Container-Apps

> Create an Ubuntu sandbox and run uname -a
							

Python SDK

📎 Quickstart

sandbox = client.begin_create_sandbox(
    disk="ubuntu").result()
result = sandbox.exec("uname -a")
print(result.stdout)
sandbox.delete()
							

TypeScript SDK

📎 Quickstart

const poller = groupClient.sandboxes.beginCreate({
  sourcesRef: {
    diskImage: { name: "ubuntu", isPublic: true } },
});
const sandbox = await poller.pollUntilDone();
const result = await groupClient.sandboxes
  .exec(sandbox.id, { command: "uname -a" });
await groupClient.sandboxes.delete(sandbox.id);
							

Inside a sandbox

Each sandbox runs as a hardware-isolated microVM with its own Linux kernel, own virtual hardware, and memory separation enforced by CPU virtualization.

Bring your agent image

Example: Dockerfile for Python agent


FROM python:3.12-slim

RUN apt-get update \
    && apt-get install -y --no-install-recommends \
       bash git \
    && rm -rf /var/lib/apt/lists/*

WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY sandbox_agent.py .

RUN mkdir -p /workspace
WORKDIR /workspace
CMD ["sleep", "infinity"]
							

Sandbox configuration

sandbox = client.begin_create_sandbox(
    disk_id=disk.id,What it boots
    cpu="1000m", memory="2048Mi",How big it is
    auto_suspend_seconds=300,When it goes idle
    egress_policy=deny_by_default,Where it can connect
    ports=[8080],Who can reach it
    volumes=[workspace_volume],What data it keeps
).result()

Sandbox lifecycle

LIFECYCLE POLICY Create Running Stopped Idle timeout or stop command stopping · snapshotting resuming Network traffic or start command Auto-delete policy Delete

📎 Sandbox lifecycle docs

Auto suspend mode with memory or disk


sandbox = group.begin_create_sandbox(
      ...,
      auto_suspend_mode="Memory"  # or "Disk"
	).result()  
sandbox.begin_stop().result()     # also happens after auto_suspend_seconds idle
sandbox.begin_resume().result()   # or send it network traffic
					

Stopped sandboxes have no compute charges and don't count against your cores quota.

Demo: Suspend and resume a sandbox

1RunningA background counter adds 1 every second.

Sandbox f2cb684b, Running, with a Stop button. The terminal reads the counter twice, two seconds apart: 1211, then 1213.

2StoppedMemory and processes are saved. No compute charges.

The same sandbox, Stopped, with a Resume button. The terminal area says the terminal is available when the sandbox is running.

3ResumedThe same process keeps counting, not back at 1.

The same sandbox, Running again after Resume. The counter reads 1347, then 1349.

Snapshots: new sandboxes from saved state


snap = sandbox.begin_create_snapshot(
    name="first-draft").result()
sandbox.delete()

restored = group.begin_create_sandbox(
    snapshot_id=snap.id).result()
							

Volumes: storage that outlives sandboxes


group.create_volume("agent-output")

writer = group.begin_create_sandbox(
    disk_id=agent_disk.id,
    volumes=[SandboxVolume(
        volume_name="agent-output",
        mountpoint="/workspace/out")],
).result()

reader = group.begin_create_sandbox(
    disk="ubuntu",
    volumes=[SandboxVolume(
        volume_name="agent-output",
        mountpoint="/data", read_only=True)],
).result()
							

Network egress: Making outbound calls


github_rule = EgressRule(
    match=EgressRuleMatch(host="api.github.com",
                          methods=["GET"]),
    action=EgressRuleAction(type="Allow"))
policy = EgressPolicy(
    default_action="Deny",
    traffic_inspection="Full",
    rules=[github_rule])
group.begin_create_sandbox(..., egress_policy=policy)
							

Outbound calls with credential injection

Egress rule: add a token for the group's identity


group_identity_token = EgressHeaderValueRef(
    managed_identity_ref=EgressManagedIdentityRef(
        identity_type="UserAssigned",
        identity_resource_id=group_identity_id,
        resource="https://cognitiveservices.azure.com",
        format="Bearer {value}"))
model_rule = EgressRule(
    match=EgressRuleMatch(
        host="myai.cognitiveservices.azure.com"),
    action=EgressRuleAction(type="Transform", headers=[
        EgressHeader(name="Authorization",
                     value_ref=group_identity_token)]))
policy = EgressPolicy(..., rules=[model_rule, github_rule])
							

Inside the sandbox: a placeholder key


AsyncOpenAI(base_url=endpoint + "/openai/v1/",
            api_key="injected-by-egress-proxy")
							

Demo: Allowed and blocked requests

Azure portal view of sandbox 9328c0ee, labeled standalone and running. The browser terminal runs curl: GET api.github.com/zen returns 200; POST api.github.com/markdown, pypi.org, and example.com return 403; and a request to the Azure OpenAI responses endpoint with no key returns a completed model response. The Egress Network Traffic panel lists the model POST and the GitHub GET as allowed, and the GitHub POST and pypi.org as denied.

Network ingress: exposing internal ports

1Run a server bound to 0.0.0.0


sandbox.exec("nohup python3 -m http.server 8080 "
             "--bind 0.0.0.0 &")
							

2Publish the port to get a public HTTPS URL


port = sandbox.add_port(8080, anonymous=True)
print(port.url)
							

Network ingress: IP access control


sandbox.add_port(8080, anonymous=True,
    ip_access_control=PortIpAccessControl(
        default_action="Deny",
        rules=[PortIpAccessControlRule(
            name="office", action="Allow",
            priority=10,
            source_cidrs=["203.0.113.0/24"])]))
							

Network ingress: Allowed users

Part 3 of 4

Back to the swarm

Sandboxes as a tool for parallel agents

Introducing the swarm Sandboxes 101 Back to the swarm Wrapping up

Swarm architecture on Azure

Microsoft Agent Framework

Open-source SDK for building AI agents and multi-agent workflows · Python .NET Go preview

🤖 Agents

An LLM that calls tools and MCP servers, with Foundry, Azure OpenAI, OpenAI, Anthropic, and more.

Here: planner, researcher, reviewer, report writer

🔀 Workflows

Graph-based workflows that connect agents and functions through explicit paths: fan-out, fan-in, and handoffs.

Here: the swarm's parallel research waves

🧰 Harness agent

Batteries included for long, multi-step tasks: planning, todos, context compaction, file access, memory, and tool approval.

Here: the standalone sandbox agent

🔌 Integrations

Model providers, agent services, tools, context providers, middleware, evaluation, and observability.

Here: Foundry web search, OpenTelemetry tracing

The swarm as a workflow graph


builder = WorkflowBuilder(start_executor=planner)
builder = builder.add_edge(planner, collector)
for r in researchers:
    builder = builder.add_edge(planner, r)
    builder = builder.add_edge(r, collector)
builder = builder.add_edge(collector, reviewer)
builder = builder.add_edge(reviewer, planner)
wf = builder.add_edge(reviewer, report_writer).build()
							

One edge per researcher, so the branches run concurrently.

🤖 plannerresearcher_0📦 in a sandboxresearcher_1📦 in a sandboxresearcher_2📦 in a sandboxcollector🤖 reviewer🤖 report_writerfan-outfan-inplan →approvedgaps:one more wave

Planner agent


class ResearchQuestions(BaseModel):
    questions: list[str] = Field(description="4-6 standalone research questions")

PLANNER_INSTRUCTIONS = (
    "Given a broad research topic, break it into 4-6 specific, "
    "focused sub-questions that together provide a comprehensive answer. "
    "Each question is sent by itself to an isolated researcher. "
    "Researchers cannot see the original topic, other questions, "
    "or other researchers' answers. Therefore every question MUST: "
    "- Be independently answerable with no prior findings... "
    "- Include the relevant subject, location, constraints... ")

planner = Agent(client=build_chat_client(), name="planner",
                instructions=PLANNER_INSTRUCTIONS)
					

Each researcher only sees its own question, so every question has to stand on its own.

Fan-out planned questions to research nodes


# Ask the planner agent for questions, as structured output
result = await self.agent.run(payload.topic,
                              options={"response_format": ResearchQuestions})
questions = validate_questions(result.value)

# Tell the collector how many answers to expect
await ctx.send_message(
    ResearchPlan(expected_responses=len(questions), ...),
    target_id="research_collector")

# Send each question to its own researcher node
for i, question in enumerate(questions):
    await ctx.send_message(
        AgentExecutorRequest(messages=[Message("user", [question])]),
        target_id=f"researcher_{i}")
					

Send each question to a sandbox

Each researcher node passes the question directly to a sandbox and returns either the response or an error.


class ResearcherExecutor(Executor):
    @handler
    async def run(self, request: AgentExecutorRequest,
                  ctx: WorkflowContext[AgentExecutorResponse]):
        question = request.messages[-1].text
        try:
            text = await self.run_in_sandbox(question)
        except Exception as ex:
            text = json.dumps({"question": question,
                               "error": str(ex), ...})
        await ctx.send_message(AgentExecutorResponse(...))
					

Start sandbox with research agent and tools


async def run_in_sandbox(question: str) -> str:
    sandbox_id = f"agent-{index}-{uuid.uuid4().hex[:8]}"
    await sandbox_mgr.create_sandbox(sandbox_id, question)        # create
    try:
        while True:                                               # execute
            await asyncio.sleep(2)
            status = await sandbox_mgr.get_status(sandbox_id)
            if status.status == "done":
                result = await sandbox_mgr.get_result(sandbox_id) # collect
                break
    finally:
        await sandbox_mgr.delete_sandbox(sandbox_id)              # delete
    return json.dumps({"question": result.question,
                       "answer": result.answer, ...})
					

Research agent inside the sandbox

Runs in each sandbox with Foundry's hosted web search, and returns a structured finding.


class ResearchFinding(BaseModel):
    answer: str
    sources: list[str]
    confidence: float

agent = Agent(
    client=FoundryChatClient(project_client=project_client, model=model),
    name="ResearchAgent",
    instructions="Use web search to find current, factual information, "
                 "then synthesize a comprehensive answer that cites sources.",
    tools=[FoundryChatClient.get_web_search_tool()],
)
response = await agent.run(question, options={"response_format": ResearchFinding})
return response.value.model_dump()
					

Collect answers for each wave

A plain executor: waits for every planned answer, then sends one dossier to the reviewer.


class ResearchCollector(Executor):
    @handler
    async def set_plan(self, plan: ResearchPlan, ctx):
        self.plan = plan
        await self.release_if_ready(ctx)

    @handler
    async def collect_response(self, response: AgentExecutorResponse, ctx):
        self.responses.append(response)
        await self.release_if_ready(ctx)

    async def release_if_ready(self, ctx):
        if len(self.responses) >= self.plan.expected_responses:
            await ctx.send_message(ResearchDossier(findings=..., wave=...))
					

Reviewer agent: approve or research more


class ReviewDecision(BaseModel):
    status: Literal["approved", "needs_more_research"]
    rationale: str
    follow_up_questions: list[str]

result = await self.agent.run(dossier, options={"response_format": ReviewDecision})
decision = validate_review_decision(result.value)
if decision.status == "needs_more_research" and waves_remaining > 0:
    await ctx.send_message(FollowUpResearch(decision.follow_up_questions, ...),
                           target_id="planner")
else:
    await ctx.send_message(ApprovedDossier(...), target_id="report_writer")
					

Report writer agent


REPORT_WRITER_INSTRUCTIONS = (
    "Given the original topic, an evidence dossier from multiple research "
    "agents, and the reviewer's assessment, produce a single comprehensive "
    "markdown report. Include an executive summary, key findings organized "
    "by theme, material limitations, and a conclusion. Cite the supplied "
    "sources where available. Do not invent evidence or sources.")

prompt = format_dossier(dossier.topic, dossier.findings) + dossier.review_rationale
result = await self.agent.run(prompt)
await ctx.yield_output(result.text)  # the final report, streamed to the app
					

Observability with OpenTelemetry (OTel)

OTel standardizes how apps emit traces, metrics, and logs, so debugging works the same across languages and vendors.

Traces

0s ───────────────────── 27s
sandbox.create
 └─ research-agent.run
     └─ invoke_agent ResearchAgent

Operations made of spans that show how a request moves through services: timing, dependencies, and context propagation.

Metrics

operation.duration  p95 26s
token.usage         24,968
requests            140/min

Numeric measurements such as latency, request counts, error rates, token usage, or any custom app metric.

Logs

INFO  Sandbox agent-2 running
INFO  Agent 3 completed research
WARN  get_status transient error

Structured log records with a message, severity, timestamp, and contextual attributes.

OpenTelemetry GenAI semantic conventions

Standard gen_ai.* span names and attributes for agent runs, model calls, and tool calls. A real span from this swarm:

Span: invoke_agent ResearchAgent (26.6 s)

AttributeValue
gen_ai.operation.nameinvoke_agent
gen_ai.agent.nameResearchAgent
gen_ai.request.modelgpt-5.6-luna
gen_ai.tool.definitions[{"type": "web_search"}]
gen_ai.usage.input_tokens22518
gen_ai.usage.output_tokens2450

Using OpenTelemetry with Agent Framework

Agent Framework has built-in support for emitting OpenTelemetry traces. Azure Monitor handles the export to Application Insights:


from azure.monitor.opentelemetry import configure_azure_monitor
from agent_framework.observability import enable_instrumentation

configure_azure_monitor(
    connection_string=os.environ["APPLICATIONINSIGHTS_CONNECTION_STRING"])
enable_instrumentation()
					
  • The same two calls run in the orchestrator and in every sandbox.
  • For sandboxes: pass the connection string in the sandbox environment, and allow the Application Insights endpoints in its egress policy.

Carry a trace across the sandbox boundary

Orchestrator: inject


def _traceparent_env() -> dict[str, str]:
    carrier: dict[str, str] = {}
    _otel_propagate.inject(carrier)
    env: dict[str, str] = {}
    if carrier.get("traceparent"):
        env["TRACEPARENT"] = carrier["traceparent"]
    if carrier.get("tracestate"):
        env["TRACESTATE"] = carrier["tracestate"]
    return env

environment.update(_traceparent_env())
							

Sandbox: extract


from opentelemetry.propagate import extract

_PARENT_CTX = extract({
    "traceparent": os.environ.get("TRACEPARENT", ""),
    "tracestate": os.environ.get("TRACESTATE", ""),
})

tracer.start_as_current_span(
    "research-agent.run", context=_PARENT_CTX)
							

One trace: separate provisioning time from research/model time, find the slow or failed branch, and verify cleanup.

Demo: follow the distributed trace

Application Insights end-to-end transaction for one swarm run. sandbox.create (813 ms) calls begin_create_sandbox and a PUT to the sandbox management endpoint. Under it, research-agent.run (14.2 s) from inside the sandbox contains ResearchAgent invoke_agent (8.9 s) and gpt-5.6-luna chat (8.9 s), with an outgoing call to the Foundry services endpoint (7.7 s). SandboxClient.delete (3.9 s) follows.

Part 4 of 4

Wrapping up

Introducing the swarm Sandboxes 101 Back to the swarm Wrapping up

Takeaways

On-demand compute

Sub-second start, zero idle cost

Without: You pre-provision servers for the peak

With sandboxes: Sandboxes start in under a second and scale to zero

Explicit boundaries

Isolation, egress, identity, data

Without: Agents can reach anything and hold API keys

With sandboxes: Egress rules, proxy-injected credentials, volumes

Explicit lifecycle

Create, stop, resume, delete

Without: Workspaces vanish on restart or keep burning budget

With sandboxes: Suspend, resume, snapshot, and auto-delete policies

One trace

Across every sandbox

Without: Parallel failures are hard to find

With sandboxes: One distributed trace from planner to every sandbox

Scenarios for sandboxes

Use caseWhat sandboxes provide
Agent workflowsPersistent, isolated workspaces that survive across task boundaries
AI code executionSafely run LLM-generated code in isolated environments with instant startup
Platform buildingBuild on the same primitive powering Microsoft services
Burst workloadsScale from zero to thousands of sandboxes on demand
Secure multi-tenant computeStrong isolation for untrusted workloads from multiple tenants
Interactive user sessionsGive each user their own isolated compute environment

Connecting sandboxes to your systems

🔌 Connectors

  • Attach once to the sandbox group; each sandbox opts in when it's created.
  • MCP connectors give agents tools to discover. API connectors give app code REST endpoints.
  • The group's identity authorizes calls, so there are no OAuth flows or tokens in the sandbox.

⚡ Triggers Preview

  • Watch a connector event, by polling on a schedule or by webhook.
  • Run a command in a sandbox, or POST to a port on a long-lived sandbox.
  • Authenticates with a managed identity, so there are no shared secrets.

Choose the execution surface

ACA surfaceBest forLifecycle
AppsLong-running services and APIsContinuous
JobsScheduled or event-driven tasksRun to completion
Dynamic sessionsManaged code execution; the platform hides the infrastructurePool-managed, ephemeral
SandboxesProgrammable isolated compute that you controlStateful: create, stop, resume, snapshot, delete
The swarm uses an App for orchestration and Sandboxes for isolated research.
A standalone sandbox does not need an ACA environment.

Source: 📎 Sandboxes overview

Keep learning