We do have Cloud Run Functions that trigger on Cloud Storage events, as well as Cloud Pub/Sub notifications for the same. Is there a specific bit of functionality you're looking for?
Hi, Brandon from GCS here. If you're looking for all of the guarantees of a real, POSIX filesystem, you want to do fast top level directory listing for 100MM+ nested files, and POSIX permissions/owner/group and other file metadata are important to you, Gcsfuse is probably not what you're after. You might want something more like Filestore: https://cloud.google.com/filestore
Gcsfuse is a great way to mount Cloud Storage buckets and view them like they're in a filesystem. It scales quite well for all sorts of uses. However, Cloud Storage itself is a flat namespace with no built-in directory support. Listing the few top level directories of a bucket with 100MM files more or less requires scanning over your entire list of objects, which means it's not going to be very fast. Listing objects in a leaf directory will be much faster, though.
Hi, GCS engineer here. GCS offered a lot of consistency from the beginning, but we didn't have strong object listing consistency at the beginning. We got that somewhere around 2017 when we moved object metadata to Spanner. See https://cloud.google.com/blog/products/gcp/how-google-cloud-...
Last time I played with this, I read through the whole thing without at any point realizing that the code was live and could be experimented with. Don't make the same mistake I did!
Also I don't know much about WebGL and am quite curious how this works. Are they compiling these shaders locally in the browser?
Yeah, that's weird. If you're going to announce a new instant messaging system, you will of course provide some reason why people should go through the effort to switch. The front page should be brimming with the pitch. "Unlike WhatsApp, with New ICQ, you can _____." What goes in the blank? Why haven't you told me?
Tech debt is most easily measurable in repetitive operational work. Measure it. Bring a story to your leaders: "We are spending 60 hours per week doing repetitive task X. With 200 hours of work, we could eliminate this work. It would pay for itself in a month."
If your leadership declines to take you up on this, escalate. If that fails, you must choose between continuing to do the repetitive operational work as instructed or leaving.
You can compose a 4TB object with a 1 byte object, or you can compose 32 150GB objects, just so long as the destination object doesn't go over 5 terabytes.
Hi, GCS engineer here. The lower limit on composing objects is one source object, in which case you are not so much composing as you are copying with style. Zero source objects is an error. I will file a note about the docs, thanks.
Well, maybe. This is Ship of Theseus situation. The good news is that, from the perspective of the new you, everything went great, and the old you doesn't have a perspective, so that's a 100% endorsement rate.
While there is not an option to set a default cache policy on a bucket, you can specify a cache policy as part of object uploads, although I admit that this can be inconvenient if you upload from multiple clients.
In addition, caching only applies to objects that are anonymously readable. Private objects are never cached.
A human would record that as clearly being PRINCE. For their described use case, reading images of business and movie names, the presence of other, small text in the picture seems quite fair.
Google has contributed quite a bit to Boto. The GCS command line tool, gsutil, is built on top of boto. If you're talking about Boto 3, though, not so much.
Once, I noticed that a minor piece of how Google's Python system tests could be very slightly cleaner and more consistent. It would be a tiny change, it seemed very safe, and it was easily accomplishable with a short sed script, but it'd also be a backward-incompatible change across the projects of hundreds of teams and thousands of build targets. I was able to make that change with only a few commits and without needing to bother most of those teams.
These sorts of small, general, large scale cleanup commits are quite common at Google, and they're encouraged. They help keep the codebase healthy. There are special groups that review them so that all of the individual teams affected don't have to bother, and there are tools to manage the additional testing and approval requirements for such a change.
At my previous company, making such a change would have been a major undertaking. I never would have considered a refactor of that scale without a critical need. They had thousands of packages, each of which had its own repository and an incredibly complex web of build and runtime dependencies. It was a nightmare, and fiddling to find a working sets of versions of internal dependencies took up way, way too much of my time each day.
The GTD book gets into this very directly. It talks a lot about the principal borrowed from martial arts of having a "mind like water." When a pebble disturbs the water, the water responds instantly with exactly the right amount of force and then returns to stillness. GTD is big on the idea of responding to emails, ideas, projects, new work items, etc, with exactly the right amount of response and then letting the problem slip out of your mind once you're confident that the thing to be done is filed and therefore will get done.
The trick to making the whole thing work is that you need to have confidence that, once something is on the list, it will get done. That's what allows you to stop fretting about the stuff on the list. But to get that confidence, you need to regularly do the stuff on the list.
Worse, Trump has issued at least 16 ethics "waivers" to at for staffers who would be banned from serving on his staff over his ethics rules. Once you make more than a dozen or so exceptions for who you're putting on your team, it's not really a rule at all.
Very cool! I like how configurable your solution is.
Google has an example app (https://github.com/GoogleCloudPlatform/kubernetes-bigquery-p...) that demonstrates getting Pub/Sub data into BigQuery, but it just does it directly as a little Kubernetes job rather than using Dataflow. I like your more serverless solution.
In addition, the $300 is only for expenses in excess of the "Always Free" usage limits, which covers quite a bit of stuff, including 5 gigs of cloud storage, 1 micro compute engine instance, and a terabyte of BigQuery queries per month.