The fundamental premise of this paper seems flawed -- take a measure specifically designed for the nuances of how human performance on a benchmark correlates with intelligence in the real world, and then pretend as if it makes sense to judge a machine's intelligence on that same basis, when machines do best on these kinds of benchmarks in a way that falls apart when it comes to the messiness of the real world.
This paper, for example, uses the 'dual N-back test' as part of its evaluation. In humans this relates to variation in our ability to use working memory, which in humans relates to 'g'; but it seems pretty meaningless when applied to transformers -- because the task itself has nothing intrinsically to do with intelligence, and of course 'dual N-back' should be easy for transformers -- they should have complete recall over their large context window.
Human intelligence tests are designed to measure variation in human intelligence -- it's silly to take those same isolated benchmarks and pretend they mean the same thing when applied to machines. Obviously a machine doing well on an IQ test doesn't mean that it will be able to do what a high IQ person could do in the messy real world; it's a benchmark, and it's only a meaningful benchmark because in humans IQ measures are designed to correlate with long-term outcomes and abilities.
That is, in humans, performance on these isolated benchmarks is correlated with our ability to exist in the messy real-world, but for AI, that correlation doesn't exist -- because the tests weren't designed to measure 'intelligence' per se, but human intelligence in the context of human lives.
The function of news is to help a democratic citizenry be critically informed, and that this kind of statistic doesn't accomplish what it set out to do, although it's certainly interesting for its own sake. I think it's a challenge of our age to figure out how to create institutions that are wise and don't simply bend to distorting pressures (money, politics, psychology).
For example, we do want terrorism over-represented relative to old-age-deaths. However, a responsible and self-aware media would really attempt to counteract 'availability bias' -- e.g. that due to the human mind what is repeated we tend to assume is actually more prevalent. But we don't have wise institutions at the moment.
The more general problem is that it is hard to quantitatively demonstrate the ways in which media fails at fulfilling its complex societal role, because it is a qualitative failure in general, although we can poke at it's edges for sure (e.g. fearmongering language probably has gone up, as has polarization on both sides of the aisle, and the amount of information-free 'babbling and speculating' in the immediate aftermath of some event has likely gone up over time).
Yeah -- I don't get why this is front-page -- reads like LLM quasi-insight:
"Through activation, lifeless equations became living systems. The neuron was no longer a mere calculator; it was a decider - a locus of transformation where signal met significance." -- wtf
The idealized (Science 1) / realpolitik (Science 2) dichotomy is both real and at first depressing. I also did a PhD in machine learning, and became quite disillusioned after seeing how the sausage was made, and how different the process is from how I had imagined it. At the same time -- engaging in 'game change' within Science 2 (perhaps not as a PhD, but after you have some security), is I think one of science's highest moral callings. The aim is not necessarily to inch Science 2 towards an impossible Science 1, but to help science to take itself more seriously (it really is a messy social process & there are ways that social process can work better or worse towards the public good -- itself a scientific question) -- and contribute towards science 2 becoming a better (and ideally better-at-self-improving) science 2.
This is naive 'populism': There's no way to avoid 'allusion' writ large -- e.g. do you object to biblical references, or to references to particular experiences that only some people have (heartbreak, death of a father)? Sure, some communities basically 'write for themselves' in a way that becomes inaccessible to outsiders w/o a lot of work. But that's fine -- I like a McDonald's hamburger as well as really nuanced flavors (for whatever reason I like nuance in how I make oatmeal that likely few others probably appreciate). Film buffs like the nuance/allusions in that medium; etc. Your comment seems like: "The stuff I like is the best and does the most for humanity" -- I think there is indeed an argument for art that is broadly appreciable, but your comment is a form of the 'gatekeeping' you criticize -- it's gatekeeping for art that doesn't require a lot of effort (for you, and those like you) to appreciate.
> Despite the HN comments complaining about it being overwhelming and a dark reflection of how awful and distracting the internet is, clearly enough people enjoyed it to get to the front page.
Is this like a massive HN wooosh -- how can this be the top-voted comment?
From Neil Postman's 1985 "Amusing Ourselves to Death":
> “With television, we vault ourselves into a continuous, incoherent present.”
> “Spiritual devastation is more likely to come from an enemy with a smiling face.”
It's less about whether we "enjoy" the stimulation, more about what kind of people we become when we lose ourselves in this bizarre sea of superstimuli. We're like reinforcement agents creating adversarial examples for each other, drawing ourselves further out of any sort of meaningful life, into a fever dream where the most desirable job for the next generation is to be famous for being famous [1] rather than do anything for any kind of deeper purpose.
> High intake of sweetened beverages was associated with higher risk for most of the studied outcomes, for which positive linear associations were found. In contrast, a low intake of treats was associated with a higher risk of all the studied outcomes.
Not sure what to make of this -- some kind of other latent explanation (e.g. that many of those with the lowest intake of treats were on a diet due to bad health?).
From the discussion section:
> One aspect to take into consideration is however that there is a social tradition of “fika” in Sweden, where people get together with friends, relatives, or coworkers for coffee and pastries (41). Thus, one could hypothesize that the intake of treats is part of many people's everyday lives without necessarily being related with overall poor dietary or lifestyle patterns, and that it might be a marker of social life.
I don't get this line of logic -- of course software has safety implications, because people use it for things in the real world. It isn't "math' that is cleanly separable from the rest of humanity; its training data comes from humanity, and it will be used towards human goals. AI is entangled with the rest of human dealings.
Whether AI poses existential threats for us or not, I'm open to either direction, but that the experts (e.g. Hinton, LeCun) are divided is reason enough to be concerned.
But applied mathematics can have ethical impact -- e.g. the concept of whether a human should trust the output of a particular language model. So GP's idea of 'trust' not applying because an object has its basis in math seems like a false dividing line. Ultimately everything can be grounded in things such as math as far as we know, although its not useful to reason about e.g. ethics from thinking about the mathematics of neuronal behavior.
The long-term impact of this paper has confused me from a technical lens, although I get it from a political lens. I'm glad it brings up the risks from LLMs but makes technical/philosophical claims which seemed poorly supported and empirically have not held up -- imo because they chose not to engage with RLHF at all (which was deployed through GPT-3 at the time; and enables grounding + getting around 'parrotness'), and uses over-the-top language ("stochastic parrot") which seems very poorly to capture what it feels like to meaningfully engage with e.g. models like GPT-4.
Ideally, meditation would spur you on to action; that's the direct aim of engaged Buddhism [1]. But more broadly, many Buddhist schools aim to encourage a direct feeling of love for all sentient beings, which if combined with the philosophy of something like effective altruism [2] (instead of Woo), could contribute to effecting meaningful systemic change.
Also -- I believe Buddhism does not apply negative connotations to 'indignation' as opposed to raw anger, i.e. I don't think it is classified as a negative state of mind to be dissolved.
The mechanism of introducing an information bottleneck, e.g. by changing from prose to poetry and trying to recover the prose, seems similar to autoencoder techniques that are popular in machine learning.
Sure, everyone can do their own research; but most won't.
There may be some bias to fact-checking, but at least it's better than relying on the candidates to do their own fact checking (i.e. usually a hugely self-serving and distorted view of reality).
The difference between sports and debates are that the post-processing in sports isn't going to change the most important "outcome," i.e. who won.
But post-processing when it comes to debates can mean overlaying information on top of the video that identifies clear falsehoods -- undermining a candidate's ability to play fast and loose with the truth to win, knowing that there's no real penalty for doing so.
So the "winner" might emerge differently if for example, news agencies didn't publish the live video, but each agency did independent fact checking (if the video were under embargo) and then each published annotated and unannotated versions.
You could still watch the vanilla version if you wanted to, but at least there would be widespread access to factually vetted versions as well.
Aren't there some set of claims that a candidate makes that nearly everyone can agree are objectively false? What's wrong with annotating a debate with that sort of information?
But by that line of reasoning, shouldn't we never hit dead ends in AI research at all -- why has AI progress been so difficult, then? Wouldn't any field of research with many dimensions of variation never get stuck on its path towards its ultimate goals, ever?
Couldn't different objective functions be structurally more difficult than others to optimize? No matter how high-dimensional the search-space, trying to create a gaming laptop in the middle ages would have been a pretty frustrating experience.
The reduction of 'humility'/'humbleness' was across a broad sample of books, not only the self-help section, and was part of a broader study [1] describing the down-trend in many words associated with virtue.
You can indeed interpret these statistics in many ways, but you first need to know the statistics.
This paper, for example, uses the 'dual N-back test' as part of its evaluation. In humans this relates to variation in our ability to use working memory, which in humans relates to 'g'; but it seems pretty meaningless when applied to transformers -- because the task itself has nothing intrinsically to do with intelligence, and of course 'dual N-back' should be easy for transformers -- they should have complete recall over their large context window.
Human intelligence tests are designed to measure variation in human intelligence -- it's silly to take those same isolated benchmarks and pretend they mean the same thing when applied to machines. Obviously a machine doing well on an IQ test doesn't mean that it will be able to do what a high IQ person could do in the messy real world; it's a benchmark, and it's only a meaningful benchmark because in humans IQ measures are designed to correlate with long-term outcomes and abilities.
That is, in humans, performance on these isolated benchmarks is correlated with our ability to exist in the messy real-world, but for AI, that correlation doesn't exist -- because the tests weren't designed to measure 'intelligence' per se, but human intelligence in the context of human lives.