
The Best AI Isn't the One That Uses the Most Tokens
By Sam Kennedy
If you've paid any attention to AI news over the past year, you've seen the headlines. Massive data centers going up in the middle of nowhere. Power grids straining to keep up. Billions of dollars flowing into GPU infrastructure. Reports of AI companies getting subsidies just to keep the lights on. It's a lot of noise, and it's easy to walk away thinking that any serious AI product must come with a serious compute bill attached.
We hear that assumption a lot. People ask us directly.
“How many tokens does Lena actually use?”
“What does your AI compute footprint look like?”
“Should we be worried about usage costs?”
The honest answer tends to catch people off guard. “Not much”. Token usage is not that high, and overage surprises are not a concern for our customers. That's not an accident, it's the result of thoughtful design choices around how we built Lena in the first place.
The Headlines aren’t about the AI Lena Runs
Most of what's driving those headlines has little to do with the kind of AI Lena runs. It's often about training massive foundation models, or running consumer AI products that need to be ready to answer almost any question about almost any topic. That's a genuinely hard problem, and it takes an enormous amount of compute to pull off, because the model has to be prepared for basically anything a person might ask it.
Lena isn't trying to do that. It has one job: operate workplace technology. It doesn't need to write poetry or explain astrophysics. That narrower scope changes the entire cost equation.
Context Does the Heavy Lifting
Think about what happens when you ask two different people to troubleshoot a conference room. The first person knows nothing about the room. They have to start from zero. What devices are in there? Which manufacturers? What firmware versions? How is everything wired together? That's a long conversation before any real troubleshooting even begins.
Now think about an engineer who has managed that exact room for three years. They don't ask any of those questions. They already know the environment, so they go straight to solving the problem.
That's the difference context makes, and it applies just as much to AI as it does to people. The more relevant context a model already has, the less work it has to do to get to a useful answer. Lena is built around that idea. Rather than asking a language model to figure out the environment from scratch every time, we hand it the specific information it needs for the task in front of it. It's not “training” in the formal sense, but functionally, that's what it feels like: Lena already knows the room.
A Quick Example
Say a user reports that a conference room camera isn't working. A generic AI assistant would need to ask a string of questions first. What camera is installed? Is it wired over USB or the network? What conferencing platform is running? Has the firmware been updated recently? Are there known issues with that model? All of that has to happen before troubleshooting even starts, and every question costs tokens.
Lena skips that entire discovery phase. It already knows which camera is installed, how it's connected, what firmware it's on, whether the configuration has drifted, and what's already been tried. It goes straight to reasoning about the actual problem instead of reasoning about what the problem might be. Quicker conversation, fewer tokens, faster answer.
Bigger Prompts Aren't Better Prompts
There's a common assumption in AI circles that more context always produces better answers. In practice, relevant context is often more valuable than simply providing more information. When a model gets exactly the information it needs, it doesn't need pages of unrelated data to sort through. It just needs what's relevant to the problem at hand.
It's a lot like GPS. Ask for directions across the country and there's a lot to calculate. Ask for directions to Conference Room B when you're already standing in the building, and the problem shrinks dramatically. Lena treats every interaction like the second case, because most of the time, that's exactly what it is.
We'd Rather Call an API Than Guess
One principle we hold onto pretty firmly is this: “We don't use AI when conventional software can do the job better”. If Lena needs a device's firmware version, it calls the manufacturer's API. It doesn't ask a model to guess. If it needs to reboot something, it sends the command directly. If it needs to check a configuration value, it compares it against a known expected value.
We reserve the AI for the parts of the job that actually require judgment: understanding what someone is asking for, reasoning through a set of symptoms, choosing between troubleshooting paths, explaining what happened in plain language, or deciding between a few possible next steps within the guardrails an organization has set. Everything else runs on ordinary software, and that architecture is a big part of why token consumption stays as low as it does.
Not Every Workflow Needs a Language Model
A lot of people assume every AI product routes every action through a large language model. Plenty of AI tools work that way, where every request goes to the model, and the model figures out what data it needs, which APIs to call, how to interpret the results, and what to say back. It's a simpler way to build a product, but it means the model is involved in almost everything, and more reasoning means more tokens and more cost.
We built Lena the other way around. If a device stops responding to a health check and an organization's policy allows an automatic reboot, that can happen through deterministic logic and a device API, no AI conversation required. Checking whether firmware matches an approved version, validating a configuration, or pulling telemetry are, in most cases, straightforward software operations. The model gets involved when there's real interpretation to do: when symptoms don't point to one obvious cause, when a few issues need to be correlated, or when the platform needs to explain what happened and recommend a next step within an organization's guardrails.
Why Our Fair Usage Policy Doesn't Worry Us
People sometimes ask why NetSpeek can afford to offer a fair usage policy that's this generous. The answer is pretty straightforward. Because Lena is architected the way it is, typical customer usage doesn’t generate the type of AI compute burden people often associate with large-scale consumer AI applications.
Most administrative interactions are short. Most troubleshooting sessions are targeted and specific. Much of what Lena does day to day runs on APIs and deterministic logic rather than language model inference. So the volume of tokens an organization would actually need to run into before hitting a real ceiling is high, and our own compute and power costs stay modest even at scale. We're not subsidizing usage to make a good headline. The efficiency is baked into how the product works.
Bigger Isn't the Right Metric
There's a tendency across the AI industry to treat scale as the whole story. More GPUs, more parameters, more tokens, more compute. For enterprise operations, none of that is really the point. Nobody buys an AI platform because they want to maximize token usage. They buy it because they want problems solved quickly, accurately, and predictably.
Sometimes the smartest AI system isn't the one burning through the most compute. It's the one that knows exactly what problem it's solving and doesn't waste effort figuring out what it already knows. That's been our approach with Lena from day one. It wasn't built to answer every question in the world. It was built to understand enterprise workplace technology, lean on context instead of guesswork, use deterministic APIs wherever it can, and bring in AI only where it actually adds value.
Which is, in a roundabout way, exactly why its token utilization stays as low as it does.
To learn more or to schedule a demo:
Visit: www.netspeek.ai
Contact: lena@netspeek.com
Demo: Book a live walkthrough
Subscribe for Updates
Stay current with the latest from NetSpeek. Learn more about new capabilities, product announcements, technical insights, and thought leadership in AI for AV.



