Axelera AI has unveiled Europa, its new artificial intelligence processing unit (AIPU) architecture aimed at physical and enterprise AI inference workloads. The Dutch company is positioning the new hardware as an alternative to GPUs for organizations that need to run models persistently on their own servers, with particular attention to power consumption, data sovereignty, and infrastructure costs. The launch arrives alongside an ecosystem of manufacturers that includes validated systems from Dell Technologies and Supermicro.
Europa from Axelera AI: the key points in 20 seconds
- Europa arrives as an AIPU for enterprise inference, physical AI, vision, and generative models.
- Axelera claims up to 6 times more tokens per second per watt than GPU-based alternatives.
- It’s offered as a standalone chip and in Edge 232p and Server 250p PCIe cards.
- The Dell XE5 and Supermicro 111AD already appear as validated systems.
- Axelera is also introducing AxScript, a Python-based language for programming its accelerators directly.
The pitch stems from a shift in AI infrastructure needs. Training large models continues to consume enormous amounts of resources, but inference — that is, running those models to respond to user requests or process data — can become a permanent workload inside companies, factories, hospitals, vehicles, security systems, and other environments.
In those scenarios, continuously paying for cloud capacity doesn’t always fit cost, latency, privacy, or data sovereignty requirements.
Axelera wants to occupy exactly that space.
The company claims Europa can deliver up to 6 times more tokens per second per watt than reference GPU solutions. This is a comparison Axelera made using internal results from its Edge 232p card against published data from competing platforms, so it shouldn’t be read as an independent result that applies to any model or configuration.
Europa takes Axelera’s architecture from the edge to the data center
Until now, Axelera AI has built much of its pitch around edge inference: processing data close to where it’s generated.
Europa extends that idea toward servers and data centers.
The architecture is available in three formats. First, as standalone silicon, so manufacturers can design their own boards. Second, as a low-profile, half-length PCIe card in the form of the Axelera Edge 232p. Finally, as a full-size PCIe card in the form of the Axelera Server 250p.
Using conventional PCIe cards serves a practical goal: adding inference capacity to existing servers without having to redesign the infrastructure from scratch.
The strategy makes particular sense for workloads that don’t need model training but instead need to run inference over long periods.
The scenarios Axelera mentions include agentic systems, vision-language models (VLMs), generative AI, and computer vision.
Also mentioned are sectors where local processing may carry additional privacy or regulatory compliance requirements, such as financial services, healthcare, defense, public administration, and legal services.
This is where one of Europa’s central arguments comes in: controlling where AI runs also means controlling where the data lives, how much it costs to process, and what infrastructure is used.
That doesn’t mean running AI locally is always cheaper than using the cloud. Total cost depends on the model, accelerator utilization, server, electricity, cooling, software, and operations. But it does offer a different cost structure and can keep certain workloads from having to continuously go out to an external provider.
The performance numbers need context
Axelera publishes four performance comparisons for its Edge 232p against a competing GPU platform, the kind of alternative to GPUs that other AI inference chip makers like d-Matrix are also chasing.
In tokens per second at batch 1, the published results are 42.5 versus 41.3 for Llama 3.1 8B; 83.4 versus 61 for Qwen3 30B A3B MoE; 11 versus 13.2 for dense Qwen3 32B; and 36.9 versus 26.3 for Llama 3.2 11B VLM.
The picture changes when energy efficiency is measured.
In tokens per second per watt, Axelera publishes:
| Model | Axelera Edge 232p | Competing platform |
|---|---|---|
| Llama 3.1 8B | 1.37 | 0.32 |
| Qwen3 30B MoE | 2.69 | 0.47 |
| Dense Qwen3 32B | 0.35 | 0.10 |
| Llama 3.2 11B VLM | 1.19 | 0.20 |
The company notes these are internal Edge 232p results compared against publicly available information from competing platforms.
There’s also an important caveat: an AI accelerator’s performance can’t be reduced to a single tokens-per-second figure.
The model, quantization, batch size, context length, precision, software, memory, and specific optimizations all play a role. An architecture that stands out on one workload may deliver different results on another.
That’s why the claim of up to 6 times more energy efficiency is interesting as product positioning, but it will need independent, apples-to-apples benchmarks to determine how Europa performs against different GPUs and accelerators in real enterprise workloads.
AxScript tries to solve one of the problems with accelerators
The other major announcement from Axelera isn’t in the silicon — it’s in the software.
The company has introduced AxeleraScript, or AxScript, a Python-based domain-specific language that allows kernel-level programming on top of its architecture.
The goal is to tackle one of the common problems with specialized accelerators: model support.
A general-purpose GPU offers enormous flexibility because it can run a wide variety of operations through different frameworks and libraries. On a specialized accelerator, the manufacturer can optimize certain workloads in exchange for imposing more restrictions on which operations or architectures are supported.
Axelera aims to lower that barrier by letting developers build custom networks, their own operators, or new architectures directly for Europa.
The company says AxScript makes it possible to work with CNNs, LLMs, VLMs, and other architectures, while retaining access to the silicon’s performance.
This matters especially because the pace at which new models appear can outstrip an accelerator maker’s ability to provide specific optimizations for each one.
According to Axelera, one engineer was able to port the Qwen3.5 model in a single day using AxScript with AI assistance.
That figure comes from the company itself and isn’t, on its own, independent proof of how long it takes to port any given model. But it does illustrate the kind of development workflow Axelera wants to achieve: software programmable enough that support for new models doesn’t depend solely on the manufacturer.
Dell and Supermicro already have validated systems
Europa also isn’t arriving on the market solely as a standalone component.
Axelera is announcing an ecosystem made up of server manufacturers, integrators, and technology providers.
The validated systems include the Dell XE5 and the Supermicro 111AD, both set up to use the Edge 232p. Systems from HPE, Lenovo, 2CRSi, Advantech, Axiomtek, and SECO also appear within the ecosystem of validated, available platforms.
This part matters for an infrastructure technology.
A company doesn’t typically buy an AI accelerator just to drop it into a server on an experimental basis. It needs to verify compatibility with the chassis, power supply, cooling, BIOS, operating system, drivers, and management tools.
The broader the list of compatible platforms, the less need there is to build dedicated infrastructure around the accelerator.
Axelera also says it has more than 600 customers deployed globally and a sales pipeline exceeding $1.5 billion. Both figures are the company’s own claims and don’t equate to recognized revenue or firm orders for that amount.
Energy efficiency is the real battleground
The energy argument comes up repeatedly throughout the launch.
That’s no coincidence.
AI is driving up data centers’ electricity demand, and organizations are starting to pay closer attention to the cost of running inference continuously.
For certain applications, cutting the power draw of each inference call can matter more than achieving the highest possible raw performance.
A system that processes fewer tokens per second but consumes much less power can be attractive when it stays running around the clock.
The tokens/s/watt metric is precisely an attempt to combine both factors.
Axelera builds its pitch around that ratio, especially for workloads that need to stay within specific power and cooling limits.
The same reasoning is even more obvious at the edge. There, it isn’t always possible to install large cooling systems or power accelerators with hundreds of extra watts.
Europa tries to carry that philosophy over to enterprise servers.
Europa also fits into Europe’s technology sovereignty strategy
The architecture’s name is no accident.
Axelera AI is a European semiconductor company, and several of its partners present Europa as a piece of a European AI infrastructure with greater control over hardware, data, and operations.
Dell, for example, has said it’s working with Axelera AI and integrator E4 Computer Engineering on AI infrastructure for Europe’s AI Factories in Italy and Luxembourg. The collaboration is framed around performance, control, and data sovereignty.
Other partners include OVHcloud, OpenNebula, 2CRSi, Telefónica Tech, and various other European and international manufacturers and integrators.
OpenNebula, which has separately been building sovereign cloud infrastructure in Europe, plans to integrate Europa into an open cloud platform to offer AI accelerators as a service, while OVHcloud highlights the combination of efficient inference with sovereignty-oriented infrastructure. These are assessments from the companies themselves, not independent certification that Europa is currently an equivalent alternative to the market’s leading GPUs.
The European angle also has an industrial dimension.
A considerable share of the world’s AI infrastructure currently depends on accelerators designed by U.S. companies and manufactured through Asian supply chains. Having European alternatives doesn’t eliminate that dependence, but it adds another option for certain deployments.
From physical AI to enterprise AI
Europa represents a change of scale for Axelera.
The company was founded with a clear focus on Physical AI, where cameras, robots, vehicles, drones, and sensors need to run inference close to the data source.
Now it’s trying to enter a much larger market: enterprise servers and data centers.
Competition there is considerably tougher.
Good performance per watt isn’t enough on its own. Compatibility with frameworks, development tools, models, and memory also matters, as do server integration, support, software availability, and the ability to keep accelerators up to date as models change.
That’s precisely why AxScript could be as important as Europa itself.
If developers can quickly adapt new models and operators without waiting for the manufacturer to add specific support, the accelerator gains flexibility.
And if that flexibility is combined with lower power draw on certain workloads, Europa would have a differentiated pitch compared with the traditional strategy of buying ever more powerful GPUs.
The launch, then, isn’t simply about presenting another AI chip.
Axelera is trying to build a complete alternative built around inference — from the silicon to PCIe cards, validated servers, compilation tools, and custom programming.
The company says Europa is already shipping and that its availability is expanding through that partner ecosystem. The coming months will show how well its efficiency numbers and its ability to run diverse models hold up outside internal testing.
For companies that want to run AI permanently within their own facilities, that comparison could end up mattering more than any peak performance figure.
Source: axelera.ai

