> You could make the argument that this well-understood process could be broken out into its own class/package/module and tested with its own public interface, but if there really is only one consumer then that's kind of a strange trade-off to make in many cases.
That's how I develop in general: a "component" does not exist because it has multiple-clients, but because it is a conceptual piece of logic that makes sense to document and test in isolation. It allows to define what is the public API of this component and what isn't. This is how software scales and stays maintainable over time IMO.
There is something to be said about individual productivity (whatever that means in a very innovative/creative environment) vs team/company output, just today I saw this in my feed: https://flocrivello.com/changing-my-mind-on-remote-about-bei...
And that's coming from someone who actually tried to build a business out of remote work (TeamFlow was the product).
I can be much more productive at home when it is about my individual contribution (me coding to deliver something unambiguous), but xxx individuals doing this does not necessarily align into a great product: that does not scale.
They claim they are writing the actual kernel code (as in the implementation of a matmul) with it, and it was presented as a "system programming language": this goes far beyond "high-level tasks" it seems.
It depends what you mean by "new subsystem" and "transitioning to": what seems like a given is that the notion of "one size fits all" of LLVM IR is behind us and the need to multi-level IR is embraced.
LLVM IR is evolving to accommodate this better, within reason (that is: it stay organized around a pretty well defined core instruction set and type system), and MLIR is just the fully extensible framework beyond this.
It is to be seen if anyone would have the appetite to port LLVM IR (and the LLVM framework) to be a dialect, I think there are challenges for this.
TensorFlow is also a runtime, yet we model its dataflow graph (the input to the runtime) as a dialect, same for ONNX. TensorRT isn't that different actually.
All of Google TPU is powered by the XLA compiler, so any MLPerf benchmark result from Google comes powered by XLA.
Anything JAX is also built on top of XLA, so you can take JAX performance as a point of comparison as well if you'd like.
The movement of paddling has a natural rotation of the shaft when you raised the fixed hand for a stroke on the other side, it's quite straightforward to figure out sitting and mimicing the movement.
During this movement if the blade aren't feathered at all you have to compensate with some bending of the wrist. The amount of rotation of the shaft induced depends on how much you raise the hand/elbow, and so is fairly dependent on your style of stroke. This is the main way I think should be approached feathering: how much vertical do you intend to paddle? From there the angle should follow to optimize for the least amount of wrist twisting.
In general paddling very vertical will come with more angle in between the blades. I practice slalom and use to have 70-80 degrees crossing, but I tend to paddle less vertically now (aging? Lack of training?) and I'm down to 60 degrees comfortably now.
But these are options, it's not a big deal to me that compiler offers special options for special use cases.
It's not clear to me if you are saying that the *default* for clang and GCC differs, aren't they both using `fno-wrapv` by default?
I don't think compiler implementations are responsible for the standard to refuse to endorse 2-complement (which is the root cause of signed overflow being UB originally if I understand correctly).
At least for GCC/Clang this isn't what O0 means. Excerpt from the GCC manual:
> Most optimizations are completely disabled at -O0 or if an -O level is not set on the command line, even if individual optimization flags are specified.
And
> Optimize debugging experience. -Og should be the optimization level of choice for the standard edit-compile-debug cycle, offering a reasonable level of optimization while maintaining fast compilation and a good debugging experience. It is a better choice than -O0 for producing debuggable code because some compiler passes that collect debug information are disabled at -O0.
(Og isn't really implemented in clang yet, but the mindset is the same)
> Don't compilers already have ways to mark variables and dereferences in a way to say 'I really want access to this value happen'?
The standard defines it, it's `volatile` I believe.
But it does not help with the examples above as far as I understand (removing the log, removing the early return, time-travel...).
> From what i can observe over years apple has almost exactly same perf culture as google and any other similarly sized company in the US
Right, if you look from very far away and put "all large US company" in the same bag.
Otherwise, if you zoom on "Silicon Valley Tech Companies" then Apple and Google's perf process and associated incentives look quite radically different in many aspects.
> The problem in this case is a chicken-and-egg problem: it's hard to get money without an education, and it's hard to get an education without money.
> For-profit education cannot solve the problem, because for-profit education is the problem.
Have you seen school that only gets paid after you start working (and based on a percentage of your salary), for example: https://www.holbertonschool.com
I like the concept in that these school are somehow "investing" in the student: they only get as successful as the student is.
Python isn't really driving the compute intensive part of ML actually, whether it's JAX, PyTorch, or TensorFlow the code is really mostly native. Convolution are implemented by hand in highly optimized libraries (Intel MKL-DNN, Nvidia cuDNN) and the Python glue is really just a light "dispatcher".
A lot of it is also asynchronous for performance: the Python code just enqueues more work to a queue which some native C++ code processes.
For TensorFlow the Python code traces an entire computation graph that is stored a protobuf and then executed by a C++ native stack, potentially remotely/distributed. Serving ML with TensorFlow does not involve any Python code in many scenarios.
Python is still quite useful for scientist to quickly glue everything together, and to describe their dataset, or when they collect result and need to produce graphs or other data analyses.
There is very little requirement on equipment to fly in the US, you don't even need a radio in the majority of the space (only when you approach towered airports and other busy / special areas). So we're far from requiring a camera :)
In case you haven't tried it yet, Pythran is an interesting one to play with: https://pythran.readthedocs.io
Also, not compiling to C but to native code still would be Mojo: https://www.modular.com/max/mojo