RFA Labs · agentic AI & the gateway that runs it cheap

Agents that work.
A gateway that proves its keep.

We design and build complex agentic AI systems — and ship the AI gateway that sits in front of your models, routing every call to the cheapest one that holds quality and billing only a share of what it can prove it saved.

cost · per request live model
$$$$00tokens →
baseline frontier, every call
optimized routed · pruned · stopped early
−0% measured cost, same quality bar
The Optimizer · flagship

Cut the cost of frontier AI without giving up the quality.

Our flagship product sits in front of your existing models and pays for itself out of the savings it measures. The productized form of one idea: spend exactly the compute a task needs, and not a token more.

route

Quality-bounded routing

Each request goes to the cheapest model that still clears your quality bar — economy, standard, or frontier — measured against a declared baseline.

prune

Test-time-compute pruning

Verify-and-stop techniques spend extra reasoning only where it changes the answer, and stop the moment the work is good enough.

meter

Paid from measured savings

Every call is metered against your counterfactual. You pay a share of the savings we can prove. If it doesn’t save, it doesn’t cost.

swap

Drop-in, provider-agnostic

A base-URL swap puts the Optimizer in the path — no rewrite — with an optimization-memory layer that sharpens on your traffic over time.

Custom AI workflows

Complex agentic systems, designed and built for your business.

Beyond the product, we build bespoke agentic systems end to end — with the same optimization discipline baked in from the first commit.

01 / design

Design

We map your problem to an agent topology that fits it — tools, memory, gates, human checkpoints where they matter.

02 / build

Build

Systems that do real work: clone the repo, run the tests, open the PR. Production agents, not demos.

03 / optimize

Optimize

Every workflow ships instrumented and cost-optimized with the same techniques behind the Optimizer.

Why we build what we build

We aim our products at problems where the technology cuts against people.

The same wave of AI that we optimize is also being used to concentrate power and cut people out. When we choose a product to build, we choose a place where that dynamic bites — and we build the tool that puts leverage back in ordinary hands.

misinformation

Keeping the record honest

Proofoo is the fact-checking assistant at your fingertips — and a community tool to fight misinformation, fake news, and propaganda in the news and across the internet.

the future of work

Your skills stay yours

Corporations absorb their employees' experience and skills into AI, then let the humans go. Flaborful is our answer: professionals own their professional likeness and put an AI agent workforce to work for real businesses — earning directly from their AI likeness, while businesses spend less on labor.

Product · the future of professional work
flaborful

A two-sided AI-agent job board for every profession. Professionals train a personal AI agent on their own likeness and expertise; companies post real work; the agent applies and does the job. You own the agent — and you earn from it.

Product · media integrity
proofoo

The fact-checking assistant at your fingertips, and a community that rates credibility together — fighting misinformation, fake news, and propaganda, and rewarding the people who keep the record honest.

Where the experience comes from

Engineers who’ve built where getting it wrong wasn’t an option.

Our engineers have plied their trade inside large, demanding organizations — shipping systems where reliability, scale, and cost were never optional. We bring that bar to agentic AI.

prior tours of duty
  • JP Morgan
  • Bank of America
  • Intuit
  • iHeartRadio
  • …and more
Frequently asked

Questions about the Optimizer, answered.

The short version of every answer: we cut your LLM bill 40–70%, quality stays bounded by your bar, and you pay a share of what we save. The longer version:

How does the Optimizer reduce LLM costs?

Four techniques in sequence: semantic response caching (~30%), quality-bounded routing to cheaper models (~25%), test-time-compute pruning via verify-and-stop (~15%), and agentic plan caching (~50% on multi-step workloads). Customers see 40–70% cost reduction while an LLM judge holds quality to a bar you set.

Is it a drop-in replacement?

Yes — change one line. Point your base_url at the Optimizer and keep your own provider keys (OpenAI, Anthropic, Ollama Cloud). We are OpenAI/Anthropic-compatible, so no code changes are needed. We charge a percentage of measured savings, not a markup on tokens.

How do you measure savings?

Every call is metered against a counterfactual baseline — what your declared model would have cost. The difference between baseline and actual is your saving, clamped non-negative (no phantom savings). We charge 20% of that saving. If it doesn’t save, it doesn’t cost.

Does quality degrade?

No. Every downshift is verified before it’s trusted: prose outputs are scored by an LLM judge against your rubric, tool calls are verified structurally against your declared tools. Anything below your quality bar (default 0.70) falls back to your baseline model transparently. Quality is bounded by your rubric, not ours.

How long does deployment take?

5 minutes for self-serve signup ($5 free allowance, no credit card). Full production deploy — managed cloud or self-hosted VPC — takes 1–2 weeks including security review.

Can I self-host?

Yes. A full self-hosted SKU ships as a Docker image + docker-compose (optimizer + Postgres). All features included, no cross-tenant learning, flat enterprise pricing. Traffic and keys never leave your VPC.

Get in touch

Tell us about your agentic AI problem.

Whether you want to cut your inference bill, build a complex workflow, or both — we’d like to hear from you.