Nord Security cuts coding-agent LLM costs by 46% with nexos.aiNord Security is one of Europe's leading cybersecurity companies, protecting millions of users across multiple products (NordPass, NordVPN). As AI coding agents became part of engineering workflows, LLM usage grew rapidly, making it harder to understand where costs came from and how to optimize them. That's when the nexos.ai AI Gateway stepped in.
VytautasBy Vytautas Vaitkevicius
46%

Reduction in total LLM costs

77%

Prompt cache hit rate, up from 14%

0

Workflow changes required

hero nordsecurity nexos ai cost

The challenge: AI costs grew faster than visibility and control

Engineering teams at Nord Security leaned on AI coding agents like Claude Code, Roo Code, OpenCode, and Zed to accelerate software development. Adoption took off fast, and that was a win. But as AI requests climbed, so did the bill, faster than the team could see, control and optimize where the money was going.

Growing LLM spend

AI adoption increased faster than infrastructure efficiency could keep pace, and costs rose with it.

Redundant token processing

Repeated prompt context was being processed from scratch instead of being reused, increasing LLM costs unnecessarily.

Limited cost controls

Spotting optimization opportunities before costs escalated was difficult.

Limited AI observability

It was hard to track where AI usage and costs originated across teams and tools.

"Our engineers adopted AI coding agents quickly, and that was exactly what we wanted. But our AI bill grew faster than our ability to understand where the money was going. We needed better visibility and cost control without slowing developers down."

Marijus Briedis

Marijus Briedis

CTO @NordVPN

The Solution

nexos.ai turned Nord Security's rising,
hard-to-see AI spend into immediate,
measurable savings. We didn't touch a single prompt, model, or agent.
Instead, every coding-agent request kept flowing exactly as before, while
the nexos.ai AI Gateway quietly reused repeated prompt context before requests reached the LLM.
All the savings, with zero disruption to how engineers work, built in.

Nord Security x nexos.ai success in numbers

Don't hear it from us. The results speak for themselves.
46%Reduction in total LLM costs
77%Prompt cache hit rate (up from 14%)
0Workflow changes required

How nexos.ai stepped in

Here's how nexos.ai solved each challenge.

A real-traffic audit that pinpointed the overspend
nexos.ai analyzed the production AI traffic already flowing through the AI Gateway to measure how effectively repeated prompt context was being reused through caching.The audit found:
  • Prompt reuse was extremely high across coding agents.

  • Caching captured only a fraction of it, with an initial cache hit rate of just 14%.

For coding agents, 70 to 80% is considered excellent, so most repeated context was still being processed from scratch on every call. The savings were sitting in plain sight.
Gateway caching at the infrastructure layer
Following the assessment, nexos.ai implemented gateway-level caching that automatically reuses repeated prompt context before requests ever reach the LLM.By processing only the new information and serving repeated context from cache, the Gateway reduced repeated LLM processing, and it stayed completely transparent to developers.
Zero disruption to engineering
The optimization ran entirely behind the scenes. No prompt re-engineering. No model migration. No application changes.
  • Same models

  • Same agents

  • Same workflow

Engineers continued working with the same tools, models, and agents while the optimization happened transparently at the gateway layer.
Better visibility and control over AI spend
Beyond caching, the engagement gave Nord Security greater visibility into production AI traffic.With a centralized gateway layer, the team could better understand AI usage patterns, identify optimization opportunities, and manage LLM spend as adoption continued to grow.

The Result

With the nexos.ai AI Gateway, Nord Security cut total LLM costs by 46% and lifted its prompt cache hit rate from 14% to 77%, into the "excellent" band for coding agents. All of it landed with no application changes, no workflow disruption, and immediate infrastructure optimization.Today, Nord Security's engineers keep moving at full speed on the same models, prompts, and agents, while leadership finally has the visibility and cost control that scaling AI requires.

"nexos.ai helped us reduce coding-agent LLM costs by 46% without changing our prompts, models, or developer workflows. The optimization was transparent to our engineers, while giving us much better visibility into AI usage and spend."

Marijus Briedis

Marijus Briedis

CTO @NordVPN

banner gradient single blob
Bring your AI spend under control

Explore more customer stories

payhawk heroPayhawk cuts security investigation time by 80% with nexos.ai
hostinger heroHostinger Horizons scales no-code AI web builder with nexos.ai
pigult heroPigu.lt integrates AI at the core of e-commerce with nexos.ai
oxylabs heroOxylabs doubles AI output and slashes costs with nexos.ai