I was already impressed by how fast 3.5 Flash was. But I've never compared it to other models in its class for coding.
Why? Coz models in that class are not very useful to me. Time saved waiting for responses usually just turns into time wasted replying to low quality responses.
Google need to release a Pro model ASAP. I am skeptical of the "maybe they don't have the compute to run it" thing. Anthropic were (probably) in that situation with Mythos and they announced it anyway - that's the obvious play for investor relations as well as hype for your product.
Coz you want to know what flooring, doors, cupboards, bath, sinks, railings, windows, etc they are putting in.
I went to view a flat this week, it was a building site. That's still the most important bit coz you get a feel for the size and shape of the space which is what really matters. But I'm glad to have the AI renders too.
Using AI for these pics is also not inherently deceptive though.
I live in an extremely overheated housing market where properties are usually sold/rented long before they actually get completed. I'm fine with landlords using AI in their renders to make claims about how the place will eventually look.
You also see people using AI to put furniture into the image (I assume they are also taking out the furniture that's actually there, belonging to the previous tenant, but doesn't fit their desired aesthetic). Again, nothing _inherently_ deceptive about this.
Main thing is just whether tenants are empowered to back out of the contract if they don't get what they were promised.
Anyone who e.g. uses AI to expand rooms/windows... Jail please.
Even if nobody is "cheating" your particular definition of cheating, the benchmarks are _somewhere_ in the super-structural gradient descent. Models are benchmark-maximising machines at some level, so I think the benchmarks are inherently a bit useless.
This is not really surprising, benchmarking _people_ doesn't work. You can only get a decent measure of someone's coding abilities by personally interacting with them. Given that models are basically person simulators it would be weird if benchmarks kept being useful as the simulation got more accurate.
I think what I've just said is basically just a more roundabout way of what you said: "Goodhart's law at work". It really is a law.
It's a extremely cool that ESPHome is able, just by existing and being good, to create this little industry of no-bullshit products. What an awesome project that is!
Yeah I was thinking airtightness might be the difference. My flat seems to be bizarrely hermetic (when you turn on the kitchen extractor fan, it struggles if you don't have a window open somewhere).
So maybe a few leaky cracks are enough that when you open a window you get a bit of a through-draft.
That's interesting coz I found the opposite, at my place to keep the level below 1k I usually have to open a window in the room I'm in, or use a fan.
I live on a noisy street so I don't usually want to do that, if I open a window at the back and keep internal doors open it will stay reasonable but significantly elevated.
So yeah I think the lesson here is you probably need to buy a sensor, different homes are gonna differ.
My home is quite small (probably 80m²) and has literally zero ventilation built in (even in the bathroom!). I live in Switzerland where it's traditional to actively ventilate your home twice a day. But that doesn't do anything for CO2. Also it's such a fucking waste of time lol. Looking forward to moving into a modern building.
IMO it's something where an intervention is often cheap enough that it's worth it even without great evidence.
But also bear in mind that regardless of "are we operating at max effectiveness", OSHA sets a legal limit of 5000ppm in a workplace, and that's about _safety_.
This article is talking about keeping levels below 1000 which is a very high standard IMO (still arguably justified given the studies mentioned). But if you are in a poorly ventilated home office you could easily hit 3000. At that point you are closer to "illegal in the US" than "earth's atmosphere".
So yeah even if you are unconvinced about micro-optimising your CO2 levels there's a very long established argument in favour of at least paying _some_ attention to it.
Looks like it's increased in price unfortunately but I like the idea, it's basically just what you would do as a DIY project but ready built. So you can either use it like a normal commercial product, or you can just fork the ESPHome config that's on GitHub and flash it exactly like any normal ESPHome project.
And I believe the accuracy is also not great on these cheap ones. The product in the OP's photo costs $200 where I live! And ISTR finding the sensor itself contributes a lot to this cost.
IIUC they also need fans. The one I have in my home has one that's actually integrated into the sensor unit.
Ah yeah I see. I guess the fallback there would be to split the service up into a remote encrypted storage layer that goes on the VPS and then host the actual service (with the decryption keys) locally?
But ISTR reading Immich kinda assumes the storage is on a plain local filesystem so you get perf issues if you do something clever under its feet. Could be out of date on that.
I really don't think you want E2EE for this. I host storage for family and friends, I haven't set Immich up yet (don't think I'd have space for everyone's photos) but the choice is between:
1. "Hey just so you know, I have access to everything you upload here".
2. "Do NOT lose your password or your data will be GONE FOREVER and I CANNOT get it back".
I definitely prefer 1 and I'm sure my users do too. They shouldn't upload it if they didn't trust me anyway.
In my case I follow it up with "and I might actually go digging around in your files if I need to debug something or you're wasting disk space". But I think you could also follow it up with "but I do promise not to look" and that would be valid too.
This whole thing only makes sense for people you're pretty close to.
(I do tell people not to back up their password managers on my system though).
I guess maybe for Immich specifically it would be nice to have a "vault" feature where people can upload nudes etc where they are willing to trade risk of loss for privacy on a per-photo basis.
No? Some subsets of the world has papyrus I guess but for most people during most of that time people were pressing text into clay and wax and stuff, it must have fucking sucked.
Then we got paper and pens and that was a pretty decent interim for a short period. Then about 100 years ago we got typing. Then about 20 years ago we reached a world where almost everyone is better at typing than they are at writing.
Obviously it's still important that people can write by hand, but expecting people to do it for more than a few hundred words at a time is idiotic. Would you like your clothes to be hand sewn too? That also worked for thousands of years (much longer than writing) but we stopped doing it for a very good reason.
Well, how many times in that 9 years have you written on paper for 2 hours straight? Even as a kid who did it regularly, it sucked.
Doing it now I really don't think I could deliver my intellectual best while worrying about if anyone can read my handwriting and whether I'm gonna cramp up by the end of the exam.
Pen and paper is just not a very good way to produce text.
As a $BIGCORP member I don't think this would be a great solution. I suspect there are plenty of vibe coding PR spammers that work for my company. And the admins of the GitHub org would not really care, making it easy for staff to contribute to third party projects is nowhere near their top priority (and policing the behaviour of their org members outside of org-owned repos is not in their mandate even if they wanted to).
You are not evaluating those questions, you are evaluating the probability of that two things happening, and you need to evaluate it better than the other people to win.
There are no easy questions, the difficulty is set by the skill/investment level of your competition.