Skip to content

Is Engy the Breakout Subnet Bittensor Has Been Waiting For?

After Kimi released the K3 model weights, Engy founder Ning said the team served the 2.8T-parameter model on 80 RTX 5090s, offering a striking proof point for frontier open-model inference on consumer hardware.

Table of Contents

Engy, a Bittensor inference subnet focused on frontier open-weight models, has drawn a rare burst of attention outside the immediate Bittensor ecosystem after founder Ning said the team ran Kimi K3, a 2.8 trillion-parameter open model, on a fleet of 80 RTX 5090 GPUs.

The announcement followed the release of Kimi K3's model weights on July 27. According to Ning, Engy had the full model running on day one at 20 tokens per second for a single stream, using GDDR7 gaming cards, standard Ethernet, and the official MXFP4 weights without requantizing the model.

That hardware choice is why the post traveled beyond a typical subnet update andinto the broader AI discussion. Frontier-scale inference is usually associated with datacenter GPUs, high-bandwidth memory, and specialized interconnects. Engy's post was groundbreaking.

"The most powerful open model on Earth, on the most abundant GPUs on Earth," Ning wrote. "Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it."

The post has now passed 750,000 views on X, giving Engy a level of visibility that has sparked Bittensor community conversations asking whether this could be the next breakout moment for the ecosystem.

What Engy Is Building

Engy is Bittensor subnet 53, an inference network that provides "verified inference" for frontier open models, meaning inference that confirms via cryptographic proof that the exact open model a user requested produced the output (not a cheaper or quantized substitute).

The project is operated by Hanlin AI, a team focused on model post-training and inference. Hanlin AI also operates TrajectoryRL, Bittensor subnet 11, which works on post-training smaller open models into specialized agents.

Engy's broader pitch combines a few ideas. It wants to serve large open models on lower-cost, widely available hardware, to prove that users are receiving the model they requested, and to use Bittensor's subnet structure to coordinate miners, validators, emissions, and routing around that service.

Engy is designed to look familiar to users. The service exposes an API at api.engy.ai and supports OpenAI-style calls, along with integrations for tools such as Claude Code, Cursor, Codex, and Hermes. That makes the product easier to understand for developers who already work with standard AI APIs, even if the backend coordination happens through Bittensor.

Kimi K3 gives Engy a particularly visible test case. K3 is Kimi's most capable flagship model to date, with 2.8 trillion parameters, native visual understanding, a 1 million-token context window, and an architecture built around Kimi Delta Attention and Attention Residuals. It's the first open-source model in the 3-trillion-parameter class, designed for long-horizon coding, knowledge work, and reasoning.

Serving a model in that class is both a performance challenge and a trust challenge.

Why The Announcement Shook up AI Talks

The viral part of Engy's Kimi K3 update was not just that the model ran, but how Engy said it ran.

Ning described the deployment as 80 RTX 5090s across 10 eight-GPU nodes. The setup used GDDR7 memory and 25GbE Ethernet rather than HBM-equipped datacenter accelerators and specialized interconnects. According to the brief, the fleet provided 2.56 TB of total VRAM, comparable in capacity to a 32-H100 configuration, and delivered 143 TB/s of aggregate memory bandwidth.

Ning framed the result as "a first for open weights," emphasizing that Engy used the official MXFP4 weights, "nothing requantized," on gaming cards and plain Ethernet.

RTX 5090s are not universally better than H100s, H200s, B200s, or other datacenter GPUs, and they are not direct substitutes across every workload, reliability profile, power envelope, or production environment. The more important impact here is that software optimization and distributed serving can push consumer hardware into workloads many people assumed required scarce, high-end infrastructure.

That matters because hardware access remains one of the largest constraints in AI. Labs, startups, universities, independent developers, and open-source communities may be able to obtain consumer GPUs more readily than fleets of HBM-heavy datacenter cards. If systems like Engy can make those GPUs useful for frontier open-model inference, the practical reach of open weights expands, helping democratize frontier model access globally.

Ning also pointed to Engy's recent work on GLM-5.2, saying the team improved performance on the same fleet from 30 to 110 tokens per second in the prior week. Engy described the 20 tokens-per-second figure as a first-day, untuned deployment, leaving room for optimization as the team learns how to serve the model more efficiently.

Why We Think Engy Could Breakout Beyond Bittensor

Many Bittensor subnet updates are difficult to explain outside the ecosystem. They often require readers to understand subnet registration, validator incentives, emissions, alpha tokens, and the mechanics of decentralized scoring before the actual product becomes clear.

Engy's Kimi K3 moment is different. A 2.8 trillion-parameter open model running on consumer GPUs is an AI infrastructure story before it is a crypto story. Add in the mix that Engy's alpha token is up ~300% in the past month, and that the decentralized/open AI narrative has made a comeback, and we have a combination on our hands that makes the Engy (mind you, still only recently launched), one of the most compelling Bittensor subnets for a broader audience.

This is, without a doubt, the subnet to watch in the coming weeks.

Engy alpha token chart

With Engy, AI developers can see the value of cheaper inference, researchers the importance of running frontier open weights outside a small number of centralized providers, and infrastructure teams the appeal of abundant hardware. Crypto-native readers can then map those practical outputs back onto Bittensor's incentive layer.

As a complete picture, Engy has combined a technically ambitious result with a message that is easy to understand outside Bittensor: frontier open-model inference on widely available GPUs, backed by proofs that users received output from the model they requested.

And that's exactly what the network needs more of. Subnets whose value is obvious before emissions, staking, or subnet economics enter the conversation.

Let's see what happens...

Comments

Latest