I started experimenting 2-weeks ago on using LLMs in a pseudo-deterministic way. I kept getting results that proved my hypothesis, which is that LLMs could be harnessed deterministically, but I could not prove why, so I kept going.
I may now have proven why. If you start your prompt input with many compiled JS binaries, it will force the LLM to take an abstract logical reasoning path that we have not seen before. I have run this thousands times against Llama-4-Maverick-17B-128E-Instruct-FP8 and Gemini-3-Flash with consistently working results.
For example, when I uploaded all Facebook binaries (i.e., FB-Static folder when loading facebook.com) at the start of my prompt, then provided my code and abstract brief, Llama-4-Maverick-17B-128E-Instruct-FP8 was able to render a fully contextual working view, considering client attributes, at a cost of of 1200 compute tokens (given 380,000 prompt input tokens).
The punchline: LLMs that we know as "math" models, significantly outperformed LLMs that we know as "abstract reasoning" models, at a small fraction of compute cost. And this may only be the beginning of the punchline.
Seeing is believing. All detailed on the link, including examples you can click and try for yourself: https://terminalvalue.net/
My experience is that it's much better than the alternative of games running natively on your phone. I would say better than the switch too in most cases for me (more battery life, shorter load times) though may depend on the game.
I also noticed a significant improvement running it on WiFi after I upgraded to a mesh network. I've only tried over cloud a few times, can appreciate connection requirements may very high, but even being able to load it up wirelessly in different parts of the house or backyard is a big plus for me.
There's no way you're going to get top quality like those sites at a place like Fiverr either. They've spent millions on branding and marketing.
I can see Midjourney already replacing the low end. The question is how good can it get, and then as the bar is raised what can be done to differentiate. Answer to the latter ironically may be going back to old school human interfaces.
I wouldn't say useless. He'll look to buy other smart peripherals that fit into the Alexa ecosystem, not a competitor's. It helps build up the brand. There's also a lot of data there to collect.
Maybe not the most lucrative revenue stream, but not nothing.
Yep, this is a feature, not just for tracking but also containment when navigating to external links. Big reason why all of those apps and others aggressively push users from web to mobile.
I've been using it. It's gotten a lot better than when it first came out. Super annoying that they released a product that was far inferior (and probably still is inferior) to Google Play music and forced everyone to switch.
Heh, SA was a huge part of my formative years, spent a ton of time (years / many many hours) lurking there. Sad to see the spiral that happened with Lowtax. Feeling very melancholic...
I can't game on PC anymore because of RSI when using mouse. Even for cross platform games, the console version usually has far superior controller support.
It's a rubber stamp like many other things in this development model.
Not saying their QA is useful, but it's not tough to imagine their QA being many layers removed from decision making. ie: Program Manager -> Product Owner -> Project Manager -> Business Analyst -> QA -- not even mentioning reporting lines.
That person in QA at the end of the chain is likely following a script written & approved by the layers above them to a T. It's highly unlikely that they are empowered to go off script or raise tangential issues. Doing so may even cost them their job.
Seems high, but this is par for the course for this type of work. You would usually estimate this in terms of months for a team. Let's say a team of 4 for simplicity -- that brings it to ~3 months.
I'd also be careful with the term "programming hours". I'm not sure how the news article got that or who said that initially, but it seems like a misrepresentation of the type of work needed. That estimate almost certainly includes everything involved in getting the code to production. You can imagine that means a lot of QA, red tape and holding the code's hand through environments.
Which car and what did you drive before? And can you quantify how much benefit you see over adaptive cruise / lane assist / etc that is standard on many cars these days?
I agree with your sentiment, just unsure what your point of reference is and how much impact the fancy AI actually has. My car (2017 model) has the features I mentioned from the factory and is great on long trips too.
> If I break something, it usually means it wasn't built strong enough to begin with. I've broken a lot of stuff over the years...
Based on my experience, this statement scares me :)
Not to pass any judgment on your impact or abilities, however the types of devs that have been the most challenging for me to work with are those with this attitude that aren't quite as good as they think they are. It can be incredibly toxic to the rest of the dev team and generally bad for business.
You need to have a very strong handle of both the business side an tech side to do this type of work effectively. Meaning: no matter how much technical debt there may be, some stuff cannot afford to be broken. Judging risk there is quite challenging as you need a holistic view. I would strongly caution people from diving in and making sweeping changes if they don't have this.
The other internal flag that went off is refactors that improve parts of the codebase in isolation while leaving a less cohesive / congruent codebase a whole. This is often worse in the long run than just patching it and actually can make changes harder.
Disclaimer: I am in mostly a management role now so you can take the above with an appropriately sized grain of salt.
Great idea. Could you explain what differentiates it from Tock? https://www.exploretock.com/ -- not that there isn't room for two solutions here, but just curious.
We've used Tock a few times in the way you described (recreate resto exp at home) and it's been awesome.
> In order to have a net positive effect, such a program would need to be able to generate an income, not necessarily a engineer market-level one (because you could have students working on this) but at least something above the minimum wage[1], not just tips.
The pay scale is going to be tough. Looking quickly through the top bounties, none of them are really student level. Issues that are truly good for beginners or someone unfamiliar with the codebase in any moderately popular open source project usually get scooped up immediately for resume building.
I can see this as being useful for abandoned projects, or as a way to promote something new. It's a much better signal that the ticket is actually meaningful and impactful than a GitHub tag.
Whether that is enough to get past psychological barriers behind transactions, I don't know. I do like the "buy the author a coffee or two" mentality of support, which I think does remove some of the transaction aspect of it. Not sure if this is what Rysolv was going for -- it feels a bit more formal.
Point well taken about culture and great examples. I completely agree with your premise.
However, tone is an extremely important part of cross-cultural training, and the tone of this article takes its premise from: _you may have a conflict of interest with group B_ to something between _group B is out to get you_ and _group B is evil_. I personally wouldn't recommend it to anyone because of this.
> As a footnote, quotation marks suggest you're quoting someone. Your claimed quote doesn't exist in the source article. Your point would be stronger if you commented on what the author wrote than your (somewhat inaccurate) read-between-the-lines.
The original post has always had the exact paragraph that I derived my subtext from in it. Emphatic quotes are a bad habit of mine, so I removed those, but c'mon -- this is a pretty common grammar mistake and the sentence preceding it pretty clearly "implies" (emphatic quotes) that it's only my interpretation :)
> That's another place cultures differ a lot: how things are implied and subtexts. People misread subtexts, which I think you did here.
Can you please explain how you would otherwise interpret the exact paragraph that I quoted? I'm usually pretty good at seeing the other side, but I genuinely can't see this one. Maybe it's because I've been in management too long (or maybe I'm too Canadian). Either way, I'm genuinely curious.
I'll quote the paragraph again directly from the article for you:
> Expert gaslighters, they are. What really makes me wonder is how the people keep doing these jobs. Many of them are in the very classes that get abused by other people regularly. How can you honestly keep doing that job when you are just enabling the abusers?
The only other interpretation I can think of is the author was referring to the people in the previous example. However, she switches from singular ("person A" / "person B") tense in the example to plural ("gaslighters" / "abusers") tense in this paragraph, so it's either an uncharacteristic grammar mistake or a generalization applied to a broader group.
The generalization is that HR people are "gaslighters" and managers are "abusers".
With that in mind, how would you interpret this sentence? "What really makes me wonder is how the people keep doing these jobs."
I can only see: _how can these people live with themselves_. How is that subtext not correct here?
I may now have proven why. If you start your prompt input with many compiled JS binaries, it will force the LLM to take an abstract logical reasoning path that we have not seen before. I have run this thousands times against Llama-4-Maverick-17B-128E-Instruct-FP8 and Gemini-3-Flash with consistently working results.
For example, when I uploaded all Facebook binaries (i.e., FB-Static folder when loading facebook.com) at the start of my prompt, then provided my code and abstract brief, Llama-4-Maverick-17B-128E-Instruct-FP8 was able to render a fully contextual working view, considering client attributes, at a cost of of 1200 compute tokens (given 380,000 prompt input tokens).
The punchline: LLMs that we know as "math" models, significantly outperformed LLMs that we know as "abstract reasoning" models, at a small fraction of compute cost. And this may only be the beginning of the punchline.
Seeing is believing. All detailed on the link, including examples you can click and try for yourself: https://terminalvalue.net/