We are on a mission to build world’s first generative media platform for developers. We are running inference on tens of thousands of GPUs, and looking for people (in all functions) to help us scale it to hundreds of thousands.
One interesting feature that gets enabled with open weights is adding new capabilities (tasks) to these editing models. They generalize quite well with low samples (30 ish). We talk about it here https://blog.fal.ai/announcing-flux-1-kontext-dev-inference-...
the main question is going to be software stack. NVIDIA is already shipping NVFP4 kernels and perf is looking good. It took a really long time after MI300X's that the FP8 kernels were OK (not even good, compared to almost perfect FP8 support in NVIDIA side of things).
I will doubt that they will be able to reach %60-70 of the FLOPs in majority of the workloads (unless they hand craft and tune a specific GEMM kernel for their benchmark shape). But would be happy to be proven wrong, and go buy a bunch of them
fal | Growth Engineer | San Francisco (on site 5 days/wk)
Help us scale generative‑media infra: hack demos in the AM, pitch partners over coffee.
You’ll build quick client libs & microsites, run data A/Bs, write content that drives sign‑ups, and hand‑hold new devs.
Need: Python, JS/React/Next.js, SQL; speed, ownership, love for gen‑AI.
Get: strong salary + equity, platinum health, unlimited “build‑something” stipend and most importantly a seat at a rocketship.
GH200 is nowhere near $343,000 number. You can get a single server order around 45k (with inception discount). If you are buying bulk, it goes down to sub-30k ish. This comes with a H100's performance and insane amount of high bandwith memory.
For traditional LLMs this might be true (especially large MoEs at bs=1) but I highly disagree with "multi-modal models" phrase since most of the models that output in other modalities are generally compute bound. Which means less flops will make the experience so much worse (imagine waiting a couple minutes for an image and hours for videos).
For anyone that wants to test the original (non-distilled) HunyuanVideo (which is an amazing model) we have 580p version taking under a minute and 720p version taking around 2.5-3 minutes in our playground: https://fal.ai/models/fal-ai/hunyuan-video (it requires github login & and is pay-per-use but new accounts get some free credits).
Hunyuan at other providers like fal.ai is cheaper than SORA for the same resolution (720p 5 seconds gets you ~15 videos for $20 vs almost 50 videos at fal). It is slower than SORA (~3 minutes for a 720p video) but faster than replicate's hunyuan (by 6-7x for the same settings).
It excels at particularly text and scene composition, as well as being able to generate vector graphics. You can use it through their website or through fal.ai https://fal.ai/models/fal-ai/recraft-v3/playground.
all high end "gaming" rigs are either using ~16 real cores or 8:24 performance/efficiency cores these days. threadripper/other HEDT options are not particularly good at gaming due to (relatively) lower clock speed / inter-CCD latencies.
FLUX.1 [dev] (non-commercial, weights-available): https://huggingface.co/black-forest-labs/FLUX.1-dev (Mindtown has a commercial license so all the images generated through it can be used for whatever purpose)
We are on a mission to build world’s first generative media platform for developers. We are running inference on tens of thousands of GPUs, and looking for people (in all functions) to help us scale it to hundreds of thousands.
Featured Roles:
- Distributed Systems Engineer, https://fal.ai/careers/4009192009
- Virtualization Engineer, https://fal.ai/careers/4146037009
- ML Performance & Systems Engineer, https://fal.ai/careers/4009191009
Remaining: https://fal.ai/careers