Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer, ML Engineer, Member of Technical Staff | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity
Artificial Analysis is an independent AI benchmarking and insights provider. We benchmark AI to help engineers and companies understand AI and make informed decisions regarding which AI technologies to use. We are fast growing with a team of 20 and have backing from investors including Nat Friedman, Daniel Gross & Andrew Ng.
We are hiring for four roles:
1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred. Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.
2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models. Strong analytical skills and proficiency in Python required.
3. Member of Technical Staff: You will develop benchmarks to test AI models and work to translate technical insights to analysis that helps companies navigate AI.
Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience (including anything you've built). Add | HackerNews to email subject line.
Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer, ML Engineer, Member of Technical Staff | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity
Artificial Analysis is an independent AI benchmarking and insights provider. We benchmark AI to help engineers and companies understand AI and make informed decisions regarding which AI technologies to use. We are fast growing with a team of 20 and have backing from investors including Nat Friedman, Daniel Gross & Andrew Ng.
We are hiring for four roles:
1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred.
Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.
2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models. Strong analytical skills and proficiency in Python required.
3. Member of Technical Staff: You will design evaluations of different AI models and work to translate technical insights to analysis that helps companies navigate AI.
4. Product Manager (AI Media Generation): You'll work closely with us to develop and enhance our media generation (Image, Video, Speech, Music) arenas and leaderboards, contributing to product strategy and execution at the intersection of creativity and AI.
Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience (including anything you've built). Add | HackerNews to email subject line.
Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer & ML Engineer | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity
Artificial Analysis is an independent AI benchmarking and insights provider. Our benchmarks help engineers and companies understand AI and make informed decisions on AI technologies.
We are hiring for two roles:
1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred.
2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models.
Strong analytical skills and proficiency in Python required.
Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience.
Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer & ML Engineer | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity
Artificial Analysis is an independent AI benchmarking and insights provider. Our benchmarks help engineers and companies understand AI and make informed decisions on AI technologies and providers.
We are hiring for two roles:
1. Full Stack Engineer:
Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred.
2. ML Engineer:
ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models.
Strong analytical skills and proficiency in Python required.
Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience.
This page has up to date information of all models and providers: https://artificialanalysis.ai/leaderboards/providers
We also on other pages cover Speech to Text, Text to Speech, Text to Image, Text to Video.
Note I'm one of the creators of Artificial Analysis.
Artificial Analysis | https://artificialanalysis.ai | Senior Full Stack Software Engineer | Onsite in San Francisco | Competitive salary + Equity
Artificial Analysis is an independent benchmarking, evaluation and insights provider for AI. Our benchmarks let engineers and companies make the best decisions on which technologies and providers to use, empowering them to build the next generation of AI applications.
We are looking for a full stack developer to support us in our analysis of AI and presenting it to the world at https://artificialanalysis.ai/.
Strong analytical skills and proficiency in Typescript & Python required. Familiarity with LLMs and AI scaling laws preferred.
Artificial Analysis | https://artificialanalysis.ai | Senior AI Analyst & Senior Software Developers | Remote or Hybrid (USA, San Francisco preferred, or Australia) | Competitive salary + Equity
We're seeking a Senior AI Research Analyst to support with benchmarking and evaluation of AI. Role involves analyzing AI systems, visualizing data, and supporting people in understanding the capabilities of AI (across different modalities).
Strong analytical skills, AI/ML research experience, and proficiency in Python/data analysis required (Typescript a nice to have). Familiarity with LLMs and AI scaling laws preferred.
We also have pricing, long/medium/short prompt lengths (decode time can vary between providers) & parallel query benchmarking + model details (ctx window, etc)
As part of our benchmarking of Groq we have asked Groq regarding quantization and they have assured us they are running models at full FP-16. It's a good point and important to check.
Groq's API performance reaches close to this level of performance as well. We've benchmarked performance over time and >400 tokens/s has sustained - can see here https://artificialanalysis.ai/models/mixtral-8x7b-instruct (bottom of page for over time view)
Hey com2kid - if you're still there, we did end up adding boxplots to show variance. Can be seen on the models page https://artificialanalysis.ai/models and on each models page where you view hosts by clicking one of the models. They are toward the end of the page under 'Detailed performance metrics'
Definitely agree with your point on Claude Instant though. Much less than half the price, much higher throughput/speed for a relatively small quality decrease (varied by how 'quality' is measured, use-case)
We have Claude Instant on the models page: https://artificialanalysis.ai/models
Can add it via the select at the top right of each card where it says '9 Selected' (below the highlight charts)
It's a combination of different quality metrics which have Perplexity, overall, not performing as well. That being said, I think we are in the very early stages of model quality scoring/ranking - and (for closed sourced models) we are seeing frequent changes. Will be interesting to see how measures evolve / model ranks change
Thanks for the feedback and glad it is useful! Yes, agree might better representative of future use.
I think a view of variance would be a good idea, currently just shown in over-time views - maybe a histogram of response times or a box and whisker.
We have a newsletter subscribe form on the website or twitter (https://twitter.com/ArtificialAnlys) if you want to follow future updates
Hi HN, Thanks for checking this out! Goal with this project is to provide objective benchmarks and analysis of LLM AI models and API hosting providers to compare which to use in your next (or current) project. Benchmark comparisons include quality, price, technical performance (e.g. throughput, latency).
We have this (and other more detailed metrics) on the models page https://artificialanalysis.ai/models if you scroll down and for individual hosts if you click into a model (nav or click one of the model bars/bubbles) :)
There are some interesting views of throughput vs. latency whereby some models are slower to the first chunk but faster for subsequent chunks and vice versa, and so suit different use cases (e.g. if just want a true/false vs. more detailed model responses)
We are hiring for four roles:
1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred. Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.
2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models. Strong analytical skills and proficiency in Python required.
3. Member of Technical Staff: You will develop benchmarks to test AI models and work to translate technical insights to analysis that helps companies navigate AI.
Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience (including anything you've built). Add | HackerNews to email subject line.