Featherless AI has put into production in Europe a dedicated cluster of NVIDIA DGX B300 servers to expand the capacity of its inference platform, which offers single-API access to more than 40,000 open-weight artificial intelligence models. The infrastructure has been deployed at a Tier III data center operated by České Radiokomunikace (CRA) in the Czech Republic, with Era Compute acting as intermediary and project coordinator.
The Featherless AI deployment in 30 seconds
- Featherless AI requested dedicated capacity based on NVIDIA DGX B300 in July and began using it in early August.
- Era Compute says the entire process took less than a month.
- Each DGX B300 packs eight Blackwell Ultra GPUs with 288 GB of HBM3e memory per GPU.
- The infrastructure is hosted at Tier III facilities run by Czech company CRA and is dedicated to inference workloads.
- Featherless AI currently serves more than 40,000 open models and is looking for additional capacity in Europe.
The project is representative of a shift starting to take shape around AI infrastructure. Access to the most advanced GPUs no longer depends solely on the major US cloud providers. European data centers are adding NVIDIA Blackwell systems to serve companies that need reserved capacity directly.
In this case, Featherless AI was after exactly the opposite of a shared, on-demand GPU: it needed dedicated, reserved physical capacity to run inference in production.
The request reached Era Compute in July 2026. The company searched for capacity among its partner providers and found availability at CRA. According to the case study published by Era Compute, the cluster began processing real workloads in early August, less than a month after the initial request.
Eight Blackwell Ultra GPUs and 2.3 TB of HBM3e per DGX B300
The choice of NVIDIA DGX B300 places the project among the most powerful commercially available GPU deployments today.
Each node packs eight NVIDIA Blackwell Ultra B300 GPUs, with 288 GB of HBM3e memory per accelerator. Combined, that’s 2,304 GB — just over 2.3 TB of HBM3e memory per system — a feature that matters especially when running large models or workloads that need long context windows.
Featherless AI will use these systems for inference — that is, to run previously trained models and respond to requests from its users and applications.
The total number of nodes making up the cluster hasn’t been publicly disclosed, so the aggregate number of GPUs or total installed power can’t be reliably calculated either. Era Compute simply describes the infrastructure as a dedicated DGX B300 cluster.
The difference compared with contracting conventional GPU instances lies in resource reservation.
Featherless needs to stay ahead of its own customers’ demand and guarantee accelerator availability. Its enterprise offering also includes dedicated GPUs, including B300-based capacity.
The company uses a model hot-swapping architecture that lets it swap models on a shared GPU fleet within seconds. That system helps explain a figure that would otherwise be hard to sustain economically: having more than 40,000 models available doesn’t necessarily mean keeping all 40,000 permanently loaded in GPU memory.
The platform loads and replaces models based on demand.
Featherless AI is also an official Hugging Face inference provider and was founded by members of the team associated with the RWKV architecture. In April 2026 it announced a $20 million Series A round co-led by AMD Ventures and Airbus Ventures.
The Czech Republic gains ground as a location for AI infrastructure
The physical infrastructure sits at facilities run by České Radiokomunikace (CRA), one of the leading Czech telecommunications and broadcasting infrastructure operators.
CRA operates 655 communication towers and runs a GPU-as-a-service division with NVIDIA Blackwell systems. According to Era Compute, the Featherless deployment sits within Tier III-certified facilities.
The choice also introduces a geographic factor that’s increasingly present in AI infrastructure decisions.
A European company can run models through an API provided from another continent, rent GPUs on a public cloud, or reserve dedicated infrastructure physically hosted within Europe.
These are distinct models, both technically and commercially.
A European location can be especially attractive for organizations concerned about data residency, latency, control over infrastructure, or their own technology sovereignty policies. That doesn’t mean that physically hosting servers in Europe automatically turns any service into “sovereign AI”: sovereignty also depends on who controls the systems, the software, the data, and the operations.
The Featherless project also involves an international mix. The accelerators are NVIDIA, a US company; Featherless AI is also US-based; CRA supplies the facilities and operational capacity in the Czech Republic; while Era Compute connects supply and demand for infrastructure.
From sourcing GPUs to running them in production in under a month
One of the elements Era Compute particularly highlights is the speed of the deployment.
According to its case study, Featherless presented its capacity need during July, and the cluster began running production workloads in early August.
Era Compute attributes the under-one-month timeline to the fact that CRA was already part of its partner network and a contractual framework was already in place.
That avoided starting the provider onboarding process from scratch. After finding a configuration that matched Featherless’s requirements, the parties could move directly to the necessary paperwork and deployment planning.
Era Compute thus works as an intermediary layer specializing in GPU infrastructure.
The company keeps track of configurations, availability, pricing, and locations across different operators. When it receives a request, it tries to match it with available infrastructure within its network.
For data centers, the approach aims to solve the reverse problem. Buying high-end AI servers requires substantial investment, and having the space, power, and cooling ready doesn’t guarantee immediately finding customers willing to reserve them.
Era Compute aims to connect that capacity with companies that need GPUs under specific terms and timelines.
In CRA’s case, the company says the project turned available data center capacity into a GPU-as-a-service contract without first having to build a full commercial relationship with Featherless.
However, the deal amount, contract length, number of DGX B300 units installed, and monthly cost of the capacity haven’t been disclosed — figures needed to economically compare the project against public cloud alternatives or other specialized providers.
Nor does the deployment appear to have ended with this first installation. Era Compute says it’s already looking for additional capacity for Featherless AI among its partners.
The need fits the way the inference market is evolving. Training large models concentrates huge amounts of GPU power over set periods, but serving them afterward to thousands or millions of users turns inference into a permanent workload.
Platforms like Featherless add another layer of difficulty: rather than focusing on a handful of commercial models, they try to provide access to tens of thousands of open models from a single API.
That requires combining GPU capacity, memory, storage, and mechanisms capable of loading and unloading models quickly.
The DGX B300 brings an especially large amount of HBM3e for that kind of work. CRA provides the European physical infrastructure, and Era Compute has acted as the meeting point between demand and available capacity.
The result is also an example of a market growing between traditional hyperscalers and buying servers outright: dedicated GPU clusters, contracted as capacity and hosted in European data centers, with resources reserved for a single platform. It echoes a similar bet Dutch provider Nebul is making with its own sovereign AI cloud built on European GPU infrastructure, and it also puts Featherless’s role as an inference layer on top of open models in a similar spot to NVIDIA’s reported $12.9 billion bid for Hugging Face, the platform where Featherless is an official inference provider.
Frequently asked questions
What is Featherless AI?
Featherless AI is an inference platform that provides single-API access to more than 40,000 open-weight models. It also offers dedicated GPU capacity for enterprise customers.
What GPUs does Featherless AI’s new cluster use?
The infrastructure is based on NVIDIA DGX B300. Each node packs eight Blackwell Ultra B300 GPUs with 288 GB of HBM3e memory per accelerator.
Where are the NVIDIA DGX B300 systems installed?
The cluster is deployed at Tier III facilities operated by České Radiokomunikace (CRA) in the Czech Republic.
How many B300 GPUs does the cluster have?
The companies haven’t disclosed the total number of DGX B300 nodes installed, so the total number of GPUs in the full cluster can’t be reliably determined.
via: era.dev

