And yes, the product right now is Keras and Tensorflow database integration + the interactive database interface. The tools around are currently under stringent testing.
Thank you for very kind comment. We are now finishing a predictor, which utilizes protein propensity data for mass-scale disorder and order predictions.
The training times obviously vary on the network architecture, software and hardware. I can safely say you can process 7200+ protein sequence with average sequence length of 120 amino acids in 2h on 2 x NVIDIA Titan XP
BTW, greetings from GROMACS group in Groningen :) I happend to do my PhD in NMR and Molecular Dynamics.
Coming back to your comments about the canonical secondary structures; I couldn't agree more with you. The problem is quite simple, how are we going to convince the >90% of structural biochemistry society to simply accept the fact proteins are bloody dynamic and X-ray / eye candy structures may have quite little to do with the "real" picture at room temperature?
Cing! Thank you for very flattering comment. Obviously, this is database only paper. Please bare in mind, the vast majority, or perhaps even >95% of protein structure prediction methods deal with canonical secondary structure classes. We want to provide a coherent data set as a benchmark + source of information.
We have in "stock" a network (obviously another paper) that will aim at propensity prediction, still in trivial alpha/coil/beta phase space.
Probably rolling out of laughter. You don't need ML to "predict" properties of molecules. Not a single physicist will but ML predicted molecule properties.
I am curious to see what will happen to Tensor Flow. I hope the code will get clean up... I also hope they will eventually pay somebody to do it, as the open source option clearly generates heterogeneous nightmare.
They are for sure! Obviously, all depends on the complexity of the problem and the willingness of programmers to install tests :/
In my company, we are ridiculously pedantic about unit testing, but even with proper level of attention to detail we sometimes fail with getting 100% code coverage.
The biggest pain are log(x)/ln(x) issues in numerical optimization.
I wouldn't dare to suggest that, yet that's the route that physics and all derivatives have adopted!
We are essentially at the crossroad and obviously, programming, which develops nowadays a bit faster than theoretical physics or mathematics, pushes in one direction.