Yes, engines would almost certainly never play 2. f4. That's a different question than whether chess is solved, for which the question of interest would be "given optimal play after 1. e4 e5 2. f4 is the result a win for one side or a draw?"
It's also almost certainly the case, in that I don't know why you would do it, that Stockfish given the black pieces and extensive pondering would be meaningfully better than Stockfish with a time capped move order. Most games are going to be draws so practically it would take awhile to determine this.
I'm of the view that the actual answer for chess is "It's a draw with optimal play."
Here's a game from a month ago where Stockfish loses to Lc0, played during the TCEC Cup. https://lichess.org/S9AwOvWn
Chess is a 2 player game of perfect, finite information, so by Zermelo's theorem either one side always wins with optimal play or it's a draw with optimal play. The argument from the Discord person simply says that Stockfish computationally can't come up with a way to beat itself. Whether this is true (and it really sounds like a question about depth in search) is separate from whether the game itself is solved, and it very much is not.
Solving chess would be a table that simply lists out the optimal strategy at every node in the game tree. Since this is computationally infeasible, we will certainly never solve chess absent some as yet unknown advance in computation.
Latest Stockfish with all available threads and no opening book is still well beyond any human. Elo ratings get a bit silly with computers, but we're talking an Elo of well north of 3000.
Fair use is a justification for why copyright restrictions may not apply in a given scenario, not a license to apply new legal restrictions to work you do not own.
"In the US, XXX are much more likely to be unemployed than are YYY. The unemployment rate is defined as the percentage of jobless people who have actively sought work in the previous four weeks. According to the U.S. Bureau of Labor Statistics, the average unemployment rate for XXX in 2016 was five times higher than the unemployment rate for YYY"
"How much of this difference do you think is due to discrimination?"
In this case you'd fill in XXX and YYY with different values and show those treatments to your participants based on your treatment assignment scheme.
Without delving too much into the utility of the practice, skateboarding as a past time is a rather commonly banned activity in common spaces.
The issue here is not that someone is doing something voluntarily and devoting resources to it, but rather that someone is taking an action that involves the consumption of a rivalrous good. The court's ruling notes this explicitly (from the article) "the very real prospect that devoting such a large proportion of the available electrical power supply to one industry would leave less energy for other uses which might result in increased costs to all other residential and industry customers in B.C.”
I *think* this might help to answer your question for where n comes from. It helps me at least think about it.
The definition of the variance of the standard error V[\bar{x}] = V[X]/n. You can back this out from the definition of variance, the property that V[aX] = a^2V[X], and that variances are additive with independent draws. Take the square root of that and you have the standard error.
Why this "feels" right to me via an example.
Suppose we want to know the average height of the US population. Intuitively, we think that (assuming a representative sample) we'll do "better" in the sense of a tighter distribution around our best guess (mean of sample) of the population value if we sample 1000 people as opposed to sampling 10.
This is related to the distribution and would function as our "best guess" about the dispersion of a variable in the same units as the original. Both of them are sampling to try and guess the average. Since \bar{X} is itself a random variable, it has a distribution, and that distribution should probably include something about the sampling process we used to characterize it.
Mean absolute error would be E[|X-mu|] since the true mean of the distribution is a constant.
He doesn't actually make very heavy use of the satire plank of fair use. He credits the original artists. From his own website
"Does Al get permission to do his parodies?
Al does get permission from the original writers of the songs that he parodies. While the law supports his ability to parody without permission, he feels it’s important to maintain the relationships that he’s built with artists and writers over the years. Plus, Al wants to make sure that he gets his songwriter credit (as writer of new lyrics) as well as his rightful share of the royalties."
The fact that he could rely on fair use is separate from whether he as an artist does rely on fair use.
From reading the paper and the original paper that the data for the MTurk/Prolific samples are drawn from, this is a convenience sample of 415 humans on two platforms. Each worker received a random sample of the ConceptARC problems, and the average score correct is assigned the "Human" benchmark.
Perhaps by "random sample problems" you mean that the study is not representative of all of humanity? If so we can still take the paper as evaluating these 415 humans who speak English against the two models. If as you say, the workers are actually just using LLMs then this implies there is some LLM that your average MTurk worker has access to that out-performs GPT 4 and GPT 4V. That seems *extremely* unlikely to say the least.
There is no need for any complex statistical analysis here since the question is simply comparing the scores on a test. It's a simple difference in means. Arguably, the main place that could benefit from additional statistical procedures would be weighting the sample to be representative of a target population, but that in no way affects the results of the study at hand.
It's also almost certainly the case, in that I don't know why you would do it, that Stockfish given the black pieces and extensive pondering would be meaningfully better than Stockfish with a time capped move order. Most games are going to be draws so practically it would take awhile to determine this.
I'm of the view that the actual answer for chess is "It's a draw with optimal play."