← Back to Home

Services

Agent Pilot, Deployment, or Fleet

Three clearly-scoped tiers. Pick the one that matches where you are, or book a free scoping call and we will tell you which fits. Fixed price after scoping, no hourly retainers.

Not sure which tier fits? Tell us about the workflow. We will recommend a Pilot, a Deployment, Fleet work, or none of it.

Typically 2 to 3 weeks

Agent Pilot

One workflow, one agent, running in production against your real data. The smallest scope that proves the thing works, or proves it does not, before you commit a budget to it.

Best for

Teams that have a workflow in mind and a demo that fell over on real inputs. Also the right first purchase if you want to see how we work before scoping something larger.

What is included

  • A scoping note covering the workflow, the inputs, the success criteria, and the failure modes we expect
  • The agent built against your real data, using sanitized examples, with no demo datasets
  • An evaluation set drawn from your actual cases, with accuracy, latency, and per-run cost reported as numbers
  • Deployment to your environment with your keys, not ours
  • A human approval gate wherever the output carries liability
  • A written handoff covering how it runs, how it fails, and what to do about it

What is not included

  • Multiple workflows. A Pilot is deliberately one agent. A second one is a second scope.
  • Deep systems integration. If the agent needs to write into three of your production systems, that is an Agent Deployment.
  • Ongoing operation. We hand it off and check in at 30 days.

Example engagement

A 40-person agency spends four hours a week reconciling client reporting across three ad platforms. Sensara builds a single agent that pulls each platform, normalizes the metrics, flags the discrepancies it cannot resolve, and posts a draft summary. Evaluated on eight weeks of historical reports before anyone relies on it. Weekly reconciliation drops to under 30 minutes of review.

Typically 4 to 8 weeks

Agent Deployment

A multi-step agent wired into the systems you already run. Evaluated on held-out real cases, deployed to your infrastructure, instrumented so output drift is visible after we are gone.

Best for

Ops and engineering leads who know the workflow, know which systems it touches, and want production software rather than a prototype someone has to babysit.

What is included

  • A scoping document covering the workflow, the data, every integration point, the success metrics, and the failure modes
  • Engineering against your real data and your real APIs, with sanitized examples
  • Evaluation against a held-out set of real cases for accuracy, latency, and cost before rollout, rerun on every material change
  • Deployment to your environment (Vercel, AWS, GCP, or on-prem) with the keys and infrastructure you own
  • Retries, cost ceilings, and structured failure handling, so a bad run degrades instead of disappearing
  • Instrumentation covering usage, throughput, and output-quality drift after handoff
  • Documentation and a runbook that lets a different engineer maintain the system without us

What is not included

  • Open-ended development. The scope is fixed. Out of scope work is a new conversation, not a quiet line item.
  • Indefinite operation. We hand off ownership, check in at 30 days, and take follow-on work only when you ask for it explicitly.
  • Strategy decks. If the deliverable you want is a document, this is not the right purchase.

Example engagement

A specialty distributor receives 200 supplier price-update emails per week, each formatted differently. Sensara deploys an agent that classifies each email, extracts the price changes, validates them against the catalog, escalates the conflicts it cannot resolve, and posts one daily summary to the buying team. Manual processing drops from 6 hours per week to under 20 minutes.

Scoped per engagement

Agent Fleet

Several agents plus the orchestration, run tracking, and observability that make a fleet operable by your team instead of by us. This is the tier we build from experience running our own.

Best for

Teams past their first successful agent who now have three of them, no shared dispatch, no idea which runs failed overnight, and no way to tell whether an agent that reported success actually succeeded.

What is included

  • A dispatcher that assigns work to agents and tracks what each one is doing
  • Run registration, heartbeats while work is in progress, and mandatory exit reporting so no run rots in an unknown state
  • Verification of what each agent claims it did, checked against the artifact it claims to have produced
  • Shared evaluation harnesses so every agent in the fleet is measured the same way
  • A console or dashboard your team uses to see the fleet without reading logs
  • Documentation covering how to add the next agent without calling us

What is not included

  • Running the fleet for you. We build the thing that makes it operable, then your team operates it.
  • A generic platform. We build the orchestration your workflows need, not a product we resell to the next client.
  • Agents you have not validated yet. Fleet work assumes at least one agent already earning its keep.

Example engagement

Sensara runs its own software delivery this way. Agents across a set of machines take a written requirements document, build the application, run it on a simulator, record a walkthrough, and prepare the store submission. Every run registers with a hub, heartbeats while it works, and reports how it exited. Runs that claim success get verified against the build and the artifact before anyone believes them.

Common questions

How do you decide which tier is right for me?

During the free 30-minute scoping call we walk through one or two real workflows. If you need to prove one workflow works before committing budget, that is an Agent Pilot. If the workflow crosses systems you already run in production, that is an Agent Deployment. If you already have several agents and no way to dispatch, track, or verify them, that is Fleet work. If none of it is worth building, we will say so.

What actually counts as an agent?

A system that takes a real input from your business, runs multiple steps against your data and tools, decides what to do at each step, and produces an outcome someone acts on. If a single prompt in a chat window solves the problem, you do not need us and we will tell you that on the call.

How do you know the agent works before we depend on it?

We build an evaluation set from your actual historical cases and hold it out of development. Before rollout we report accuracy, latency, and per-run cost against that set as numbers, not impressions. After rollout the instrumentation keeps those numbers visible so you can see drift instead of discovering it from a complaint.

Is everything fixed price?

Yes. Every engagement is scoped and priced up front after the scoping call. We do not bill hourly retainers. Out of scope work is a new conversation with its own price, not a quiet line item added to the invoice.

Do you operate the agents you build for us?

No. We hand off ownership at the end of every engagement, including code, infrastructure, keys, documentation, and a runbook. We schedule a 30-day check-in to review usage and the evaluation numbers, and we take follow-on work when you ask for it, but day to day operation is your team.

What stacks do you work in?

TypeScript and Python against the major LLM APIs, deploying on Vercel, AWS, GCP, or on-prem depending on data sensitivity and the team that will maintain it. We integrate with the systems you already run rather than asking you to adopt new ones.

What about data privacy and security?

We do not retain client data after a project ends, and every deployment uses customer-owned API keys and infrastructure. Sensitivity constraints and any approval gates the work requires are surfaced during the scoping call so nothing surprises either side at the implementation stage.

Not sure which tier fits?

Book a free 30-minute scoping call. We will walk through one of your workflows and recommend a Pilot, a Deployment, Fleet work, or none of it. No deck, no pitch.

Curious how we work?

Read about our four-phase process

The AI Workflow Diagnostic, from the first call to the 30-day check-in. Same anti-hype voice; more detail on how each engagement runs.

Read the process