A search engine index is an economic exchange between the website and the publisher.
To massively (over)simplify the argument to its essence (and ignore other important points): the publisher goes through the trouble and expense of creating the content
The publisher then allows its content to be copied by a search engine only because being shown in search results gets it traffic back. The traffic it gets in return has value, and the publisher is happy for this arrangement to continue as long as the value of the traffic is more than the cost of producing and serving the content.
Brave offering a "license", for its own financial benefit, to "allow" others to use the content for LLM training gives zero benefit to the original publisher. This is why I use words like "sleazy" to describe Brave's position.
This argument applies to Google and Microsoft. Right now both are failing at citing sources in their generative AI search results. That is terrible and I hope it's fixed soon, as otherwise they're being sleazy scrapers as much as Brave is.
Finally, I wholeheartedly disagree they what Brave is doing is for the "greater good". The fact they charge extra for the "license" to use the content for LLM training shows that.
There is a a difference between a human being able to access content vs a search engine indexing it (and in the case of Brave, "licensing" it on).
I share your concern about Google having this much power, and I'd add that Microsoft Bing is equally bad but gets away with it because they're smaller. Still, the final decision about which search engine indexes a website is purely the publisher's.
This is explained more in the article I referred to, but briefly: Brave delegates crawling to normal Brave browsers, so it's a huge IP addresses pool, not a single IP address or range.
Also, these search crawls by the browser do not identify themselves beyond the Brave standard UA header, namely a plain Chrome user-agent string.
That would be bad, and it is already bad that Google and Microsoft control so much of search queries, but the decision about which search engine indexes a website is purely the publisher's.
The major problem with Brave search is their position about indexing and licensing content against the wishes of the website publisher. Their robot does not identify itself, meaning the publisher cannot use the standard robots.txt to block its crawling if the publisher so wishes. Incidentally, the robots.txt file has been used in court cases litigating if a search engine is legal or not.
Even worse, they state that Brave search won't index a page only if other search engines are not allowed to index it. It is morally not their right to make that call. A publisher should have full control to discriminate which search engine indexes the website's content. That's the very heart of why the Robots Exclusion Protocol exists, and Brave is brazenly ignoring it.
Even worse than that, the Brave search API allows you (for an extra fee) to get the content with a "license" to use the content for AI training? Who allowed them the right to distribute the content that way?
It's a much more nuanced position that can be summarized as "make sure you create good content, however you create it". A focus on quality, not process, is reasonable.
In simplified terms, did you find everything you could have possibly found? Looking at the formula in the article, it includes the false negatives, that is, items you misclassified as negatives when you should have considered them positives. And because that happened, you didn't find them in the set, that is you "forgot them". The opposite of forgetting is... recall.
Another place this idea comes up is a search engine index. If the algo doesn't find, for a given query, documents in the index it should have (falsely classified as not matching the query), it will have lower recall.
The faint lines are the cell walls and the bright spots in the middle would be the DNA. I can believe this is what they're going for with a bit of squinting.
Before anyone thinks this (and similar) approaches are a way around the GDPR's cookie consent tracking crackdown: It's not.
The GDPR talks about online identifiers, of which cookies, IP address and fingerprints are examples. If you read any regulator's guidance carefully, you'll see they talk about "cookies and similar technologies", with just "cookies" being used alone for brevity.
To rephrase tracking of any kind is the issue, not cookies. Don't mistake the implementation for the activity.
Disclosure: Founder of a non-tracking web analytics service because of this exact issue.
The privacy policy is very not suited for this service. The most important point is that you're based in Germany based on the address in the policy, but there isn't a single mention of the GDPR. That and the ePrivacy Directive are what count for you the most. My recommendation is don't use a free policy generator and get proper advice. I appreciate this isn't something commonly seen as a launch blocker, but it's important to sort it out properly.
Find your German state data protection authority, and invariably you'll find they have great guidance.
Yes, and also cookie IDs. Both are called out as examples in recital 30:
“Natural persons may be associated with online identifiers provided by their devices, applications, tools and protocols, such as internet protocol addresses, cookie identifiers or other identifiers such as radio frequency identification tags. This may leave traces which, in particular when combined with unique identifiers and other information received by the servers, may be used to create profiles of the natural persons and identify them.”
Looks good! I'm the founder of a similar service (Blockmetry). Obviously non-tracking web analytics is the future!
I'm curious why you chose to host the data yourself instead of giving customers the data immediately at the point of collection. That's the path we chose for Blockmetry as it genuinely required to be a non-tracking web analytics service and makes it impossible to profile users. Any service that hosts its data would still be open to being untrusted on the "no tracking no profiling" argument.
Thanks,
Pierre
PS - YC Startup School founders: ping me via the forums and get an extended-period free trial.
Speaking of commas, you're missing one after the end of the interrupting phrase in your last sentence (should say ", as well as the state of Maine,"). It's a pet peeve bigger than the lack of Oxford commas, and definitely affects readability and may affect meaning.
I don't have access to the raw log files from the customers, so can't give you a percentage. All I'll say confidently is that my service processes a lot of bot traffic that needs to be filtered out before reporting.
BTW, are you the same Peter Hartree on this Segment thread? https://community.segment.com/t/1889n1/how-common-is-client-... It would appear we've crossed paths before on this topic. Please do email me if you want to talk properly. That Segment thread has my email.
I operate a service that measures this (see another comment on this discussion), and all I'll say is you'll be very surprised how many bots actually execute JS, especially stealth bots. You have to be careful either way.
https://news.ycombinator.com/item?id=36993739