Hey, I’m acutely in the market (considering moving away from Google)
2 Qs:
1. How does OpenCage correctness/completeness compare to Google Maps API, especially in rural and industrial regions where you have addresses like “AcmeCo Industries, 234-XY Unit C, Jebel Ali Free Zone, Dubai”? I’d like to confidently query the most precise location that still matches/contains my query.
2. Do you support querying by business names? Google’s geocoding doesn’t return the business name in the result (that’s a separate API), but it does use business names to resolve queries.
What would happen if Russia attacked these power plants directly? Are they built to "fail safe" even when hit by a missile, or is there a big risk of nuclear meltdown?
I encourage anyone to read the (surprisingly plain-english) first few pages of the decision, but here is the gist of it:
> In January 2024, the court issued a post-trial opinion finding that the award was subject to review under the entire fairness standard [...] the defendants bore the burden of proving entire fairness, they failed to meet their burden, and the plaintiff is entitled to rescission. [...] The defendants responded by putting the rescinded compensation plan [...] to a stockholder vote for the stated purpose of 'ratifying' it. [...] The defendants then moved to 'revise' the post-trial opinion based on the stockholder vote, asking the court to flip its decision.
> The motion to revise is denied. [...] The large and talented group of defense firms got creative with the ratification argument, but their unprecedented theories go against multiple strains of settled law. [...] First, the defendants have no procedural ground for flipping the outcome of an adverse post-trial decision based on evidence they created after trial. [...] Second, common-law ratification [...] cannot be raised for the first time after the post-trial opinion. [...] Third, [...] a stockholder vote standing alone cannot ratify a conflicted-controller transaction. Fourth, [...] material misstatements in the proxy statement [defeat the ratification]. Each of these defects standing alone defeats the motion to revise.
> The fee petition is granted in part. The plaintiff’s attorneys asked for $5.6 billion in freely tradeable Tesla shares. [...] That was a bold ask. [...] Delaware courts award fees based on a percentage of the value of the benefit achieved [...] yet [...] a fee award 'can be so large that typical yardsticks [...] must yield to the greater policy concern of preventing windfalls to counsel.' [...] $5.6 billion is a windfall no matter the methodology used. [...] To reach a reasonable number, this decision [...] uses the $2.3 billion grant date fair value to value the benefit achieved. [...] Applying a conservative 15% to that figure results in a fee award of $345 million—an appropriate sum to reward a total victory.
For anyone else confused by what actually happened, here is a summary compiled from various sources around the conviction and the related 1MDB scandal:
---
Jho Low, a Malaysian financier, masterminded one of the largest embezzlement scandals in history through 1Malaysia Development Berhad (1MDB), a sovereign wealth fund intended to spur economic development. Over $4.5 billion was siphoned from the fund to finance a lavish lifestyle, high-profile investments, and extensive political influence campaigns. Fleeing justice in Malaysia, Low focused on cementing his power in the U.S., including efforts to influence the political landscape and suppress investigations into his crimes.
Pras Michel, a founding member of the hip-hop group Fugees, became entangled in Low's schemes, leading to his conviction on 10 criminal counts. Michel first met Low in 2006, and by 2012, he was a key player in Low’s efforts to use his ill-gotten wealth to influence U.S. politics. Low funneled $20 million to Michel to gain access to then-President Barack Obama’s re-election campaign. Knowing direct contributions from foreign nationals were illegal, Michel orchestrated a scheme using straw donors and political committees to route Low’s money into the campaign. Michel also used funds to buy seats at fundraising events and pressured wealthy acquaintances to participate.
By 2017, Michel’s involvement deepened as he acted on behalf of both Low and the Chinese government without registering as a foreign agent. In exchange for millions, Michel attempted to influence the Trump administration to drop the U.S. investigation into Low and to extradite Chinese dissident Miles Guo, a target of Beijing. These actions violated federal law, which requires registration for such foreign lobbying efforts.
Michel was also convicted of laundering millions of dollars tied to the 1MDB embezzlement and attempting to obstruct justice by pressuring straw donors to support his version of events during the investigation. The trial revealed Michel’s use of burner phones to contact witnesses, an act he later admitted was misguided. His defense argued that Michel was unaware of the legal boundaries and acted on bad advice from his attorney, including the use of artificial intelligence to craft his closing argument—a controversial decision.
The prosecution presented Michel as a knowing participant in a broader conspiracy to influence U.S. politics and aid foreign interests. Testimony from high-profile witnesses, including actor Leonardo DiCaprio and former Attorney General Jeff Sessions, underscored the scale of the scheme. Michel was ultimately convicted of conspiracy, campaign finance violations, acting as an unregistered foreign agent, money laundering, and witness tampering.
The frustrating thing with SOC2, or pretty much most compliance requirements, is that they are less about what’s “technically true”, and more about minimizing raised eyebrows.
It does make some sense though. People are not perfect, especially in large organizations, so there is value in just following the masses rather than doing everything your own way.
Read the notes in the link you posted. I don’t think it says what you think it says.
In May 2020, the definition of M1 (monetary supply in “cash”) was changed to include savings deposits. They changed this not due to some conspiracy, but because savings accounts were deregulated to remove withdrawal limits, effectively rendering them cash-equivalent, and thus necessary to include in M1 metrics.
I.e. the 80% spike has nothing to do with money being printed.
I mean a normal passenger on a normal plane making a normal trip to an office building and finding a hidden location where to tape a small box with an arduino in it. Maybe even on the outside so you can use solar power? Though it only needs to last long enough to compromise a machine inside the network.
This would be nothing new, I remember ages ago in the days of WEP that you could buy a small box that would collect enough handshakes to let you crack the WEP password.
You’d need to go a level below the API that most embedding services expose.
A transformer-based embedding model doesn’t just give you a vector for the entire input string, it gives you vectors for each token. These are then “pooled” together (eg averaged, or max-pooled, or other strategies) to reduce these many vectors down into a single vector.
Late chunking means changing this reduction to yield many vectors instead of just one.
In a perfect market, the market maker who sells you that option offsets it with correlated assets in the other direction, eg by buying or selling stock that is sensitive to the election.
Large trading firms exist on finding and exploiting small arbitrages between various correlated assets. If you assume a perfect market with infinitely many participants and infinite liquidity, then this “works” - there is no distortion at scale.
This is awesome, but I'm not sure what the long-term use case for the intersection of low-latency integration and non-production-stable is? I'm saying this as someone with way more experience than I'd like to in using reverse-engineered APIs as part of production products... You inevitably run into breakages, sometimes even actively hostile platforms, which will degrade user experience as users wait for your 1day window to fix their product again.
Though I suppose if you can auto-fix and retry issues within ~1minute or so it could work?
I see - I suppose that’s a fair viewpoint to have!
I’m not much of a python programmer but my experience with the language would make me tend to agree actually. There are bigger fish to fry and so the effort to go after this relatively tiny sardine is perhaps not worth it.
Why would you go through all the hassle of setting up an air-gapped system, only to stop at enforcing strict code signing for any executable delivered via USB?
1. Slack has network effects: we connect with customers on Slack (or M$ Teams for enterprise…)
2. I don’t want to innovate on “back office”. Slack works, is sufficiently affordable, and costs no social credit with employees.
3. I know I won’t run into problems in the future. This kinda ties into (2), I don’t want to innovate on back office, but to make it concrete: Deel, Rippling and the other M$ AD clones all integrate with Slack to set up permissions and SSO with zero effort.
4. Slack has lots of sensitive info about us and pur customers, which makes it SOC-2 relevant. I want to use “industry standard” tech for anything compliance related. Though I don’t recall that this would have ever been a problem on security questionnaires, and not many years ago Slack was the “young kid on the block” themselves & managed, so idk if this is actually a valid point.
—
If you give me a feature parity clone of slack for half the price, I’d certainly switch, but anything less than feature parity and I probably wouldn’t. I don’t need to or want to take risks on internal tooling.
My understanding was that the extra parameters required for the second attention mechanism are included in those 6.8B parameters (i.e. those are the total parameters of the model, not some made-up metric of would-be parameter count in a standard transformer). This makes the result doubly impressive!
Here's the bit from the paper:
> We set the number of heads h = dmodel/2d, where d is equal to the head dimension
of Transformer. So we can align the parameter counts and computational complexity.
In other words, they make up for it by having only half as many attention heads per layer.
Some of the "prior art" here is ladder networks and to some handwavy extent residual nets, both of which can be interpreted as training the model on reducing the error to its previous predictions as opposed to predicting the final result directly. I think some intuition for why it works has to do with changing the gradient descent landscape to be a bit friendlier towards learning in small baby steps, as you are now explicitly designing the network around the idea that it will start off making lots of errors in its predictions and then get better over time.
This certainly contributed to people preferring card over cash, making merchants loose ~3% per transaction.
That ship has long sailed, but it does male you wonder: if everything was priced at increments of, say, quarters, would enough people still use cash to offset the lost sales from the allegedly less appealing pricing?
I see what you're saying, but I don't think it applies in this case. Correct use of jargon helps domain experts communicate with higher precision, and papers tend to be written by domain experts for consumption by other domain experts.
Of course there are some (possibly many!) papers where jargon is abused to make something sound smarter. Sometimes this can also happen unintentionally.
In this case, "compute-optimal X" is standard terminology used in large-scale ML model design for finding the most optimal tradeoff with regards to compute when trying to achieve X.
Here, the paper is about finding the optimal model size tradeoff when training on LLM-generated synthetic data. Imagine you have a class of LLMs, from small to infinitely large. The larger the LLM, the higher the quality of your synthetic data, but you will also spend more compute to generate this data ("sampling" the data). Smaller LLMs can generate more data with the same compute budget, but at worse quality.
The paper does some experiments to find that in their case, you don't always want the largest possible LLM for synthetic data (as previously thought by many practitioners), instead you can get further by making more calls to a smaller but worse LLM.
2 Qs:
1. How does OpenCage correctness/completeness compare to Google Maps API, especially in rural and industrial regions where you have addresses like “AcmeCo Industries, 234-XY Unit C, Jebel Ali Free Zone, Dubai”? I’d like to confidently query the most precise location that still matches/contains my query.
2. Do you support querying by business names? Google’s geocoding doesn’t return the business name in the result (that’s a separate API), but it does use business names to resolve queries.