HexStellar Founder on Cutting GPU Inference Energy Use Without Touching the Model - Startup Fortune
Your subscription could not be saved. Please try again.
Your subscription has been successful.
AI Briefing
5 most important AI updates after 8pm every day, curated for you.
Subscribe free
Search
SF
Aug 16, 2026 · 9:08 PM<br>– Users OnlineOnline
US|<br>INTERNATIONAL|<br>ASIA|<br>SF Featured
Subscribe✉
Most Read
Codex Users Are Losing Banked Rate…<br>Myspace Owners Confirm They Are Planning…<br>AI Model Distillation Becomes the New…<br>BitMart announces it is shutting down…<br>Codex Users Are Losing Banked Rate…<br>Myspace Owners Confirm They Are Planning…<br>AI Model Distillation Becomes the New…<br>BitMart announces it is shutting down…
Home<br>Entrepreneur
HexStellar Founder on Cutting GPU Inference Energy Use Without Touching the Model
Brayon Pieske, founder of HexStellar and Trust Carbon Infrastructure, on how a carbon data platform produced a measured cut in GPU inference energy use, why the company publishes its most conservative numbers, and what it means for data centres that have already bought their hardware.
Amilia Bon
Aug 12, 2026 · 1:33 PM ·<br>5 min read<br>2.1K reads
this.classList.remove('sf-copied'),1500)" aria-label="Copy link">
Most efficiency work in AI asks you to give something up. Shrink the model, drop the precision, accept a slightly worse answer in exchange for a smaller bill. Brayon Pieske, founder of HexStellar and Trust Carbon Infrastructure, spent months refusing that trade, and says the result is a measured improvement in GPU inference energy efficiency that leaves the model itself untouched.
The company has published a conservative floor of its results rather than its best ones. Pieske spoke to StartupFortune about how a carbon data platform ended up producing an efficiency layer, why the measurements took months, and what it means for a data centre that has already bought its hardware.
The product came out of a question you kept being asked, not a roadmap.
We were building the infrastructure behind Trust Carbon. Last year, in 2025, people kept asking us things like how do you run vision AI on a smartphone without internet, and how does the battery last that long. After enough of those questions we decided to investigate more carefully. Because we build almost everything from scratch, we realised we had done something different. At that moment we did not fully understand what it meant. Only after deeper investigation did we see it was not just an internal fix. That is when GPU inference energy efficiency stopped being a side effect and became the product itself. We filed patents before we published anything.
What did the months between noticing it and filing actually involve?
Those months took time on purpose. When you see a number that strong, the first job is to try to kill it. We checked whether anything similar already existed, compared it against other approaches, and did extensive research. We designed a measurement protocol that could survive an audit before we allowed ourselves to believe the result. Only after it kept holding did we move to protect it. Three provisional patent applications were filed in June 2026, after the protocol and the runs.
Why measure on real hardware rather than model it?
Modelling is easy to adjust so it looks good. Measuring on the actual equipment under the same conditions, and publishing both the strong results and the cases where almost no difference appeared, is what builds credibility. What matters for users is simple: lower energy use, cooler operation, and more capacity on the same hardware. In the sealed measurement, energy per completed request went from 506.7 joules to 243.8 joules, a 51.9 percent reduction, while completed requests in the same hour more than doubled. Memory use stayed the same. Measured, not modelled.
Why Sonera Merged Compliance and Deliverability Into One Caller ID Layer<br>Sonera formed by merging Contact Center Compliance (DNC.com) and Pure CallerID to unify call compliance and deliverability into one caller ID platform.
caller ID compliance and deliverability layer<br>· why merge DNC and caller ID
You tested across several model families. What would have changed if it had only worked on one?
This was the most surprising part for us. At first we expected the effect would only appear on certain systems. When we saw it working across different architectures and platforms with the same kind of integration, that was the moment that shocked us most. It made clear this was not a trick tied to one model or one type of machine. We test across independent model families and only publish each package when it survives the same protocol, one at a time.
You publish your most conservative numbers rather than your strongest. Why?
Most companies publish their best number. We do the opposite on purpose. Our strongest results are so strong that if we led with them, many people would assume it was just marketing. So we chose to publish...