With utensor.ai, you can probably try this out today. We are currently working on integrating CMSIS-NN with uTensor.
CMSIS-NN are these MCU SIMD optimized functions.
Thanks man!
You summed it up nicely! This project is about making the trade-off between speed and cost. There is greatness in either-end of the spectrum.
You are right that the F767ZI has 512kb of RAM. However, the MLP code has been tested on boards with 256kb of RAM. Changes are on the way to pull that number even lower, significantly.