You cannot bruteforce this. Exhibiting a unknotting of K with n moves only gives you an upper bound u(K) <= n. Proving u(K) = n is an entirely different matter.
Proteins are linear molecules consisting of sequences of (mostly) 20 amino acids. You can see the list of amino acids here: https://en.wikipedia.org/wiki/Amino_acid#Table_of_standard_a.... There is a standard encoding of amino acids using single letters, A for alanine, etc. Earlier versions of ESM (I haven't read the ESM3 paper yet) uses one token per amino acid, plus a few control tokens (beginning of sequence, end of sequence, class token, mask, etc.) Earlier versions of ESM were BERT-style models focused on understanding, not GPT-style generative models.
I actually prototyped a system like this, mostly as an exercise to learn about crypto. You can't feasibly host or verify proofs on-chain, so you need external trusted verifiers (e.g. oracles). Making sure the oracles can't front-run proof submission is a challenge. Standard formal proof system (like Lean) are sufficiently expressive, although they weren't built for this and need to be modified to make sure a proof hasn't introduced any additional axioms, as you note. The proof system also becomes a point of attack, so you'd probably want multiple, independent verifiers (which themselves have been formally proved correct). I believe these exist for some proof systems, although I'm not sure about Lean's kernel.
Ultimately, I don't think this is really practical, and investing in AI proof agents is the way to go.
They have a 20B parameter model. I think the primary dataset for these open models is The Pile: https://arxiv.org/abs/2101.00027 (web scrape, pubmed, arxiv, github, wikipedia, etc. There is a nice diagram on page 2 that summarizes the contents.)
Where are you losing people in your hiring process? Are you getting no initial applications? Do you give the coding exercise but they never do it? Do you give an offer but they turn it down? There is a lot of speculation in this thread but you should have the data.
Mathematician here. What do you want this for? Even if you had them, you probably wouldn't understand the definitions anyway.
As others say, there is no standard, and conventions vary by subfield, publication, author and over time. This is esp. true at the research level, where the mathematical content is still being worked out. Subfields have certain conventions, and well-written books and papers will normally introduction notation or include an index of notation, esp. if the notation is novel or they different from the usual conventions. You could start compiling something like this by going through the standard undergrad and grad textbooks for each subject.
One of the best technologists I ever worked with denied his interest in technology until he was around your age. He was a professor at a top school in CS and started some innovative and impactful companies. I quit my job at 34 to study math and got my PhD at 40. After that, I left math to work in biology and I run a data science/engineering group at a premier biology research institute. I will probably change things up again before I'm done. I am not unique, there are many examples of this:
It is pretty clear Jim Keller did something pretty remarkable at Apple and then AMD (I know less about his work at Tesla). I tried to dig into the stuff he's said and written to understand what he did and how he did it. Say what you want about Fridman's interview style, that interview was probably the most insightful thing I found.
The Broad Institute of MIT and Harvard was launched in 2004 to improve human health by using genomics to advance our understanding of the biology and treatment of human disease, and to help lay the groundwork for a new generation of therapies.
The Hail team's mission is to build tools to enable rapid analysis and exploration of biological datasets (100s of TB and tripling yearly). We are committed to open science and everything we do is open source. We currently develop in Python, Scala/Java, and C/C++ and use Spark, Kubernetes, Google Cloud Platform (GCP) and AWS, but will use any tools we need to get the job done. Come help us build the future of big scientific data analysis.
We have two positions:
Update: The Site Reliability Engineer position has been filled.
We also have a front-end/designer position that will be posted shortly. Email below, get in touch if you're interested.
You don't need experience in biology or our particular technologies. We work in a highly multi-disciplinary environment (with software engineers, biologists, bioinformaticians, doctors, operations, statisticians, etc.) Self-improvement is a fundamental part of our culture. You must be excited to be challenged and learn new things.
I'm the hiring manager. Get in touch with me directly if you have any questions: [email protected].
We're quite a bit smaller but have similar numbers: 15-20m right now. We're dominated by build time (build caching might help) and schlepping docker images.
Lots of other jobs at various levels throughout the institute. Biology knowledge generally note required (I had none), although it helps (but be prepared to learn).