Thank you very much! Appreciate your comment and encouragement.
I conceived this idea while working on a project with multiple people to evaluate multiple LLMs. Too many API keys floating here and there in the code. So I said, what if we could use only one master key for all the APIs and hide the real key from users. Then, LlaMakey was born.
A smaller model has less capacity and thus is less prone to overfitting, which is a reason of hallucination. Overfitting means that a model cannot adjust its output based on the input that is unseen during training.