Anyone wanting to train an AI language model without directly generating CO2 had few practical options until now. The Munich data center of Deutsche Telekom offered both: computing power from renewable energy and waste heat use for the neighboring Tucherpark neighborhood. There a consortium of nine German institutions trained Soofi S from March to May 2026, an open language model with 31.6 billion parameters, available on HuggingFace since mid-July.
What Soofi S is and how it came to be
SOOFI stands for Sovereign Open Source Foundation Models. The consortium, coordinated by the AI Federal Association, comprises nine institutions: the L3S Research Center, Fraunhofer IAIS, Fraunhofer IIS, the German Research Center for Artificial Intelligence (DFKI), the University of Würzburg, TU Darmstadt, Berlin University of Applied Sciences, plus two startups. The Federal Ministry for Economy and Energy financed the project with roughly 20 million euros within the IPCEI-CIS initiative for European cloud and AI infrastructure.
Training ran from March to May 2026 on up to 512 NVIDIA B200 GPUs and consumed roughly 253,000 GPU-hours. Soofi S is a Mixture-of-Experts model with a hybrid Mamba-Transformer architecture: of 31.6 billion total parameters, roughly 3.2 billion are active per pass. This means high throughput with relatively low operating energy consumption. The model was trained on 27 trillion tokens from English and German text. On June 17, the consortium presented first performance results; since mid-July, the weights are publicly available.
What's special: training on renewable energy in Munich
What European AI projects rarely communicate is a central project goal here: the entire training ran on 100 percent renewable energy. The data center was cooled with water from the Isar river. Waste heat was fed into the neighboring Tucherpark neighborhood. By comparison: training a model the size of GPT-3 generated roughly 300 tons of CO2-equivalent according to University of Massachusetts calculations. Soofi S renders that number moot through energy source choice, though the exact savings in tons were not published.
On standard benchmarks for open models, Soofi S ranks first among European open-source systems in English and German: it surpasses Spain's Alia 40B, the European EuroLLM with 22 billion parameters, the Swiss Apertus 70B, and the U.S. reference point OLMo 3. For industrial practice, more relevant than benchmarks are tasks like analyzing technical specifications, processing regulatory documents, code generation, and agent-based systems: exactly what the consortium designed Soofi S for.
What the model cannot yet do
Since July 16, the weights on HuggingFace carry the status "preview checkpoints". The license is not yet finalized. Independent analysts at consulting firm Wavect write: Soofi S is not yet a production-ready replacement for Claude, GPT, or larger Chinese open-weight models. For complex reasoning, extensive software development, or scientific deep analysis, smaller models like Soofi S clearly fall behind proprietary flagship models.
A direct performance comparison with Claude 4, GPT-5, or Google's Gemini line on real business workflows has not yet been published. Consortium partners point to the strategic target audience: companies, administrations, and research institutions that cannot use U.S. or Chinese models for data protection or compliance reasons and need transparent, self-hostable systems. For this market, Soofi S is the first European candidate with competitive German language competence.
Why the consortium matters more than the model
Soofi S is, by its own account, the first building block of a model family. The next step: a model with roughly 100 billion parameters. The path there costs far more than 20 million euros and more computing than 512 GPUs. For comparison: Meta's Llama 3.1 with 405 billion parameters was trained on 16,000 NVIDIA H100 GPUs over multiple weeks. The U.S. and China invest billions, not millions, in their frontier models.
What the consortium offers is something else: demonstrable infrastructure. Nine institutions have proven they can jointly build a competitive model while maintaining reproducible, transparent conditions that respect European regulation and data protection. That is the foundation on which a European 100-billion-parameter model could emerge, if funding follows. Whether it follows hinges on political decisions in Berlin and Brussels, not on the technology.
