We are a young team, and the creators of the Apertus LLM, the currently leading open-data open-weights AI model.
Join us to work on cutting edge LLM training in the open. We do pretraining, alignment, reasoning, multilinguality and multimodality - all at the intersection of engineering and research.
This is a joint team between ETH Zurich and EPFL in Lausanne, running on the Alps supercomputer (one of the largest public institution GPU cluster).
Visa sponsoring possible, work language is English.
We hear you, nevertheless this is one of the very few open-weights and open-data LLMs, and the license is still very permissive (compare for example to Llama). Personally of course I'd like to remove the additional click, but the universities also have a say in this.
The pretraining (so 99% of training) is fully global, in over 1000 languages without special weighting. The posttraining (See section 4 of the paper) had also as many languages as we could get, and did upweight some languages. The posttraining can easily be customized to any other target languages
common crawl anyway respects the CCbot opt-out every time they do a crawl.
we went a step further because back in old ages (2013 is our oldest training data) LLMs did not exist, so website owners opting out today of AI crawlers might like the option to also remove their past contents.
arguments can be made either way but we tried to remain on the cautious side at this point.
we also wrote a paper on how this additional removal affects downstream performance of the LLM https://arxiv.org/abs/2504.06219 (it does so surprisingly little)
we released 81 intermediate checkpoints of the whole pretraining phase, and the code and data to reproduce. so full audit is surely possible - still it would depend on what you consider 'practical' here.
we kept all 1800+ (script/language) pairs, not only the quality filtered ones. the question if a mix of quality filtered and not languages impacts the mixing is still an open question. preliminary research (Section 4.2.7 of https://arxiv.org/abs/2502.10361 ) indicates that quality filtering can mitigate the curse of multilinguality to some degree, so facilitate cross-lingual generalization, but it has to be seen how strong this effect is on larger scale
Yes this is an interesting question. In our arxiv paper [1] we did study this for news articles, and also removed duplicates of articles (decontamination). We did not observe an impact on the downstream accuracy of the LLM, in the case of news data.
No, the model has nothing do to with Llama. We are using our own architecture, and training from scratch. Llama also does not have open training data, and is non-compliant, in contrast to this model.
We are a young team, and the creators of the Apertus LLM, the currently leading open-data open-weights AI model.
Join us to work on cutting edge LLM training in the open. We do pretraining, alignment, reasoning, multilinguality and multimodality - all at the intersection of engineering and research.
This is a joint team between ETH Zurich and EPFL in Lausanne, running on the Alps supercomputer (one of the largest public institution GPU cluster). Visa sponsoring possible, work language is English.
https://careers.epfl.ch/job/Lausanne-AI-Research-Engineers-S...