After spending an entire career doing 'by hand' (and a helluva lot of molecular orbital calculations) on the problem this post is about, i've got to tersely weigh in with: there's (still) not enough available data given the size of protein 'phase space' to hope for a proper covering with one's trained up linear algebra model. Or typed another way: you've got to include at some stage some physical modeling parameters, like molecular orbitals [1], otherwise the 'response curve' will only optimize if one gets quite lucky, (which is actually unlucky as then you'll delude yourself into thinking it's a generally applicable, which it isn't). For instance, swap in a carboxylic acid moiety where there was previously an aldehyde, a protein side-chain flips over, and you're in a completely different corner of the energetic 'galaxy'.
Had a QuantSci Prof who was fond of asking "Who can name a data collection scenario where the x data has no error?" and then taught Deming regression as a generally preferred analysis [1]
Circa 1980, as a hobbyist beekeeper with six hives nearby Seattle, they came down with foul-brood. (It never was certain if it was American Foul-brood or European) Duly reported and the county agent came in, sealed them, and carted them away. For an additional fee (which I paid) they would fumigate them and return just the hive bodies, but none of the frames, (some of which would contain the infected brood). Those were burned. I believe the fumigant at the time was phosphine[1]
Lately in advertisements there have been a lot of circular QR codes[1]. Which often apparently present data outside of the alignment patterns. Is that merely an artistic device, or is there some standard permitting information extending beyond the square region?
..and curiously, with the little speaker icon 'on', it produces no audio alarm for me, (firefox 103.0.2). Whereas this[1] one produces a quite notable sound.
As more of a story of external impressions, i'm from Seattle and I postdoc'ed in Germany and I got more than my share of the opinion on English from those who could speak English there (damn near 100% at a University). The first problem they had was placing my accent; they couldn't. Finally they concluded that i was some-sort of Canadian (which is pretty accurate). Then they did their own versions of U.S. accents. Almost universally these divided into New York and Texan. Their impression of the states, based on either vacationers or hollywood, confirmed their view that people from the U.S. were either loud with a Texas accent or loud with a New York accent. We decided that it was a natural sampling bias based on the overhearing of the loudness.
After gmail recently greeted my mutt/IMAP request (which had worked perfectly for over a decade!) with "login failed", i've embarked upon the slow winding road to getting mutt to use OAuth2, or is it OAUTHBEARER?, for gmail. Out of a dozen or so hints spanning the last couple of years (many which depend on having python2 still around), i'd humbly promote these links to best hints:
It does not require training data - instead it uses statistical information about the frequencies of the letters and n-grams in the English language.
and from this it should also be noted that it won't apparently be able to extract passwords, as least those which aren't "n-grams in the English language".
At some point in every bioinformatics lecture i always manage something akin to: "Learn awk! (or perl) You'll need it. Your data will come from various disparate sources, and you need to get them into some well-defined useful format from the get go."
A layman's speculation followed by a layman's (at least in terms of immunology) question: introduce a novel mRNA, cells produce novel protein, it is presented to the immune system as something to foment a defense again: vaccine. But the article talks about using this mechanism to replace a functional copy of a genetically missing enzyme. How is it assured that this will assume the opposite goal and not become another target for the immune system to guard against?
[1] e.g. https://proteindf.github.io/