Trying the fuse feature, seemed the most interesting:
> spaghetti ≈ trumpet
> Both spaghetti and a trumpet can be difficult to eat without making a mess—spaghetti with its long, slippery noodles, and a trumpet with its wide, flared bell.
Nvidia's enterprise GPUs are surprisingly unreliable. Working on a 128 GPU A100 cluster on AWS, 1 would fail every few days. I didn't have any insight on whether it was a hardware or software failure.
You also need the optimizer (e.g. Adam)'s state, which is usually double the parameter's size. So if using fp16, one parameter takes up 6 bytes in memory.
No it won't. Large language models are trained on 1,000 - 50,000 GPUs. No one's going to buy hundreds of Mac pros to mount them in a datacenter for training ML models.
> Ok, when people start racing to post these at midnight, and beg for upvotes on top of it, this experiment has officially jumped the shark. I'm going to bury this post and ask you all not to post any more of them.
Only one account (whoishiring) is allowed to make regular feature posts that we don't kill as duplicates. (That's for the obvious reason of preventing karma sweepstakes and race conditions.) Should we make this "Idea" thread a regular feature? I've thought about it quite a bit. I think the answer is no.
Experiments are worth trying, but this one has gone on for a month now and I don't think it has cleared the bar [1]. Something about having all these ideas in one place makes the whole less than the sum of its parts. The threads seem to me to have gotten less interesting as they've become more regular.
I'm sorry to disappoint those of you who disagree. But our job is to optimize HN for quality and I don't think the quality is high enough here. Ideas are better in the wild. Let's discuss them as they come up organically, rather than try to organize an idea-fest.
1. https://news.ycombinator.com/item?id=7682938