Frankensteining LLMs
Session Abstract
TNG’s R1T Chimera models reached over 10 billion daily tokens on OpenRouter. How can a small consultancy produce its own LLMs? In our case: by Frankensteining them. This talk motivates: you can adapt LLMs to your own needs without being a big research lab. Including theory, technicalities, and lessons learned.
Session Description
We demonstrate how with comparably little effort, often even few GPU resources, one can manage to still adapt LLMs, instead of purely being satisfied with the results of the big players. TNG is not a research lab, yet we have been successfully publishing our own models and are still working in directions how to adapt even huge LLMs with very reasonable efforts.
For Haystack, this is admittedly a long-shot, but I was recommended to submit. While we do not cover adaption of LLMs to improve search capabilities, our talk could motivate to go that direction.
Fabian Klemm
TNG Technology Consulting GmbH