CANARYONE

Your company’s AI stack, without the assembly.

Companies end up stitching together chat, knowledge, coding, routing and admin from different vendors. We’re building the box they should have come in, over a model layer we measure ourselves.

The benchmarking tool ships today · the platform is in build · every figure on this site links to the run behind it

ONE SYSTEM

AI people can use.
Infrastructure the company can control.

We’re building three surfaces on one shared layer, so a company does not have to buy them separately and then make them agree with each other.

Ask

In build

Ask questions across company information without copying documents into another tool.

Build

In build

Give engineers an OpenAI-compatible API and coding tools on the same model layer.

Control

In build

Manage people, model access, budgets, policy and audit across every surface.

One CanaryOne layer

  • Shared identity
  • Company knowledge
  • Policy & budgets
  • Measured model layer

Usually: 5 tools · 5 consoles · 5 bills CanaryOne: one layer

CANARYONE INTELLIGENCE

The model layer is measured, not guessed.

A shelf of models picked off a public leaderboard is a guess about somebody else’s workload. We read what the market advertises every day, and we measure what real work costs on each route.

Market intelligence tells us where to look. Workload benchmarking tells you what actually wins.

MARKET INTELLIGENCE

The same model, 12× the price depending on where it runs.

gpt-oss-120b, as advertised by the companies serving it. Each dot is one host.

$0.030 per million $0.350 per million

The rule marks the middle of the market. Half of the companies serving this model charge under 12 cents per million tokens.

14 models · 220 live endpoints
Read several times a day
Captured 17 August 2026 at 12:55 UTC

WORKLOAD BENCHMARKING

The cheapest tokens can cost the most per finished task.

One workload, one model, ten routes. Three of them below, on identical tasks.

Three of the 10 routes measured in run e860167a, showing attempts finished and cost per finished task.
Route Finished Cost per finished task
Route A 6 of 12 $0.0472
Route B 9 of 12 $0.0686
Route C 12 of 12 $0.1852

run e860167a · 10 routes · 4 tasks · 3 repeats each
Measured 29 July 2026 · open the full report →

Two kinds of evidence. The market figure is what a host advertises, not a measurement of how well it serves the model. The route figures are one local run of our own tool, which is why they are lettered — the size of a gap like this repeats from run to run, and the name at either end does not.

MOVE AT YOUR PACE

Planned

Move workload by workload.

Start with the models you already use. Add control and measurement first. Move a workload only when an open model proves the better choice.

Today

Keep the models you already use

  • OpenAI and Anthropic
  • existing integrations
  • no model change required

Through CanaryOne

Add control + measurement

  • one identity
  • budgets + policy
  • audit
  • workload measurement

When one wins

Move one workload

  • open-weight model
  • regional or specialised host
  • only after it proves itself

YOUR COMPANY

Planned

Your knowledge. Your region. Your rules.

The four questions that decide whether a company can put a tool like this near its own documents at all.

Connect knowledge

Read from the systems where company information already lives.

Carry permissions

Carry the identity and permissions a company already has.

Control policy & spend

Set model access, spending limits and usage in one place.

Choose where it runs

Choose the region and infrastructure the workloads run on.

WHEN YOU NEED EVIDENCE

Simple for users. Inspectable when it matters.

A buyer who has to justify the choice should be able to see the working. The measurement layer is the part of this that already exists, and every run it does writes an artefact you can open.

run e860167a · 10 routes · 4 tasks · 3 repeats each

Open example report →
A CanaryOne report: one row per route, with attempts finished, cost per finished task, judge score and latency percentiles beside each other.
Workload benchmarking
Today
Run it on your own repository and get one row per route.
Route and model evidence
Today
Every figure on this site carries its date and links to the run behind it.
Audit and policy record
Planned
Part of the admin surface we’re building.
Regional and host choice
Planned
Part of the model layer we’re building.

The report is one static file, written to <repo>/.c1/runs/<runId>/report/index.html, so it opens from disk and travels in a pull request.

Stop assembling your AI stack.

Start with the models and the tools you already use. We’re building one place for a company to use them, control them, and improve what runs underneath.

EARLY ACCESS

Get early access.

We'll reach out when early access opens.