A plot point in an episode of Person of Interest, where AGI "Samaritan" funds a charity (free tablets for students, Samaritan access pre-installed) to take over education and recruit more mercenaries.
I'm building this, for (mostly) non-scientific non-fiction works (books, articles, news, etc.). Launching soon, with about 7,500 books indexed.
Generally, what I found useful to build a graph between "topics" or entities was to use a HyDE[1] prompt to generate possible distinct definitions and then build a nearest-neighbor network from that. This successfully identifies related concepts in the truly abstract sense rather than literal entities.
The full use case includes quantisation, which the repo points out uses a large amount of system RAM. Of course that’s not required if you skip that step.
With 16 threads, about 140ms per token for 30B, 300ms per token for 65B
I should also mention that 65B should be able to run on 64GB systems. Total system memory consumption on M1 Ultra is about 67GB when running nothing else.
That's impossible to judge. LLama is a foundational model. It has received neither instructional fine tuning (davinci-3) nor RLHF (ChatGPT). It cannot be compared to these finetuned models without, well, finetuning.
Not to forget the Toshiba AC100 with the original Tegra, which was one of the first usable ARM laptops (netbooks?) and helped much of Linux desktop support development for ARM in the early days.