Build — Agentic AI development

Agents that do real work, within limits you set.

An agent that can read, decide and act is only useful if you can say what it may do, prove what it did and stop it at once. Onega designs, builds and evaluates AI agents and multi-step workflows that work in your systems, with every tool call checked by policy, consequential actions approved by a person and a record you own.

01 — What we build

From one task to a governed workflow.

Task agents

Single-purpose agents that triage, extract, look up, draft or reconcile, and hand a prepared result to a person.

Multi-step workflows

Agents that plan across several tools and systems, with checkpoints between steps and a defined way back when a step fails.

Knowledge agents

Answers grounded in your documents and records, each with its sources, and a plain “not found” when the sources are silent.

Tool and system integration

Connectors to your ERP, CRM, service-desk, document and mail systems through their APIs or the Model Context Protocol. Credentials are never given to a model.

Evaluation harnesses

Test sets drawn from your own cases, scored before release and on every change, so a new model or prompt must prove itself first.

Operations

Monitoring, cost budgets, drift checks and incident runbooks for agents in daily use.

02 — Autonomy

Autonomy you dial, not flip.

Every action class gets its own level. Agents start at L0 in a pilot, and nothing climbs a level without evidence on your own cases and your approval. A guardrail breach is designed to lower the level automatically.

LevelWhat the agent doesWhat we require first
L0 · ObserveRuns in shadow; nobody acts on its resultsA baseline and an evaluation set
L1 · SuggestProposes; a person confirms or correctsEvidence and sources with every proposal
L2 · One-clickPrepares an action set; a person executes itTyped actions and deterministic validation
L3 · Automatic with reviewRuns a qualified action; people are notified and can reverse itA reversal window and continuous monitoring
L4 · AutonomousRuns a narrow, proven action class unattendedSampling audit, an error budget and automatic downgrade

03 — Controls

Governed at every call.

When agents run through OneVeer, these controls sit in the request path. Where you already operate an AI gateway, we build to its controls and document any gap.

  • One governed gateway for every model and tool call, so policy, redaction and budgets apply to every agent.
  • A signed kill switch scoped to one agent, one model or everything, enforced before any model is reached.
  • Single-use approvals from a person for consequential tool calls.
  • Credentials held in a vault; models see results, never keys.
  • Guard models on inputs, retrieved context and outputs.
  • A record of every step (inputs, sources, tool calls, approvals and outcomes) kept in your environment.

04 — How a build runs

Two weeks to scope, six to eight to prove.

  1. 01

    Discovery Sprint · 2 weeks

    The task, the tools, the data, the risks, the autonomy level to start at and acceptance criteria on your own cases.

  2. 02

    Build · 3–5 weeks

    Agent, connectors and evaluation set built in your environment and tested on a holdout sample.

  3. 03

    Supervised run · 2–3 weeks

    Real work at L0 or L1, every proposal reviewed, effort and errors measured against the baseline.

  4. 04

    Handover

    Runbook, configuration, evaluation results, known limitations and a decision: extend, raise a level or stop.

05 — FAQ

Questions architects and CISOs ask.

Which agent frameworks do you use?

The simplest one that meets the requirement, often plain code around a model’s tool calling. Where your teams have standardised on a framework or on the Model Context Protocol, we build to it. The governance sits in the gateway, not in the framework, so the choice stays reversible.

Can agents run on local models?

Yes. OneVeer supports tool calling on local models as well as through cloud providers, and a published alias can move an agent between them without changing the agent.

What stops an agent from being manipulated by what it reads?

Everything an agent reads is treated as untrusted. Tools are allow-listed per agent, credentials never reach the model, guard models inspect inputs and retrieved context, and consequential actions need a person’s approval. Detection is fallible, so the design assumes some attacks will get through and limits what they could do.

Who is accountable for what an agent does?

Your organisation, which is why each agent has a named owner, a version, an evaluation record and operating limits agreed before it goes live, and why its record stays with you.