> For venue recommendations [...] we do not rely purely on the language model. We embed both user requirements and venues into vector representations and retrieve candidates using similarity search. Hard constraints such as capacity and dates are applied first, and results are ranked before being presented.
Huh this surprised me as a forgone opportunity.
I heard second-hand about the process for organizing our last offsite. Searching for venues was not the time-consuming part.
The time-consuming part was actually engaging with the venues to confirm specific details not available online. Our teammate who did this engaged with _hundreds_ of venues. It was a lot of work on their part ... and probably not the most fun part of their job.
Hey we're also a Vertex tuning customer in a similar spot. We're seeing other capacity issues, although not a leap in latency. Can you DM me? I'd love to trade notes. https://x.com/hellofromjames
I love Cerebras. I also love that they've started to scale rate limits to useful levels (which is relatively new).
I still don't know how long they'll support our chosen model.
On Oct 22 I got an email saying that
```
- qwen-3-coder-480b will be available until Nov 5, 2025
- qwen-3-235b-a22b-thinking-2507 will be available until Nov 14, 2025
```
That's not a lot of notice!
I don't want to spend all my time benchmarking new models for features I already built. I don't want my users' experience to be disturbed every few months.
Google[1] also has a "long context" pricing structure. OpenAI may be considering offering similar since they do not offer their priority processing SLAs[2] for context >128K.
You and your friends should email me with your resume and anything you're proud to have built. I'll extend that to any MIT senior/recent grad who wants to discuss moving to SF and helping us apply LLMs to build product features that solve interesting customer problems.
I'm at [email protected]. Include "[responding to HN thread 43614795]" in the title. I'd love to chat.
I am grateful for GCP's quotas that help us prevent similar own-goals.
While this specific error is something we know to avoid, I'm sure quotas have helped us avoid the pain of other errors. So I'm somewhat sympathetic.
I think it's important to read the language of and judgements in the post in the context of someone who just got a large unexpected bill (expensive lesson).
I noticed Anthropic updated their prices, but haven't seen this posted anywhere.
Claude Instant is now 10% of Claude 2's pricing: $0.80 per million input tokens, and $2.40 per million completion tokens (down from I think $1.63 and $5.51 respectively).
Altman mentioned[1][2] earlier that they were working on a "stateful" API for release this year.
> 2023: A stateful API — When you call the chat API today, you have to repeatedly pass through the same conversation history and pay for the same tokens again and again. In the future there will be a version of the API that remembers the conversation history.
Maybe it's an RAG-based thing, but that'd be underwhelming given the promise.
Wizard of Oz, or true magic?
(In the same interview, Altman also claimed progress toward releasing million-token context windows this year. Wowzers)
Inflation means tomorrow’s money is worth less; it impacts cashflows more when they are further out.
When the cost of money is near-zero, today’s values of near and distant cashflows are similar. When the cost is high, they are very different.
I’m not sure if you mean in your question that a project shown to track inflation will be unaffected. This is somewhat true — we see this in inflation-adjusted bonds etc. But inflation is far from a uniform effect, and I’ve never seen a pitch include inflation in its estimates…
I'd start with goals. Take some time (you have it available!) to think about what you'd like to achieve at different levels, such as:
- In your career
- In this field
- At this company
Then break them down, and then break them down again. What can you do, in bite-sizes pieces, to move towards them?
But I wouldn't stress too much. Remember back to when you were busy, and how much of that became growth. If you're like most people, it wasn't very much.
Use the time to achieve what _you_ want, but don't forget to enjoy it too :)
I really enjoyed reading the fiction-reflecting-reality novel "The Shipping Man" [0]. It's a fun read (though the story is better than the writing).
In particular the author (IIRC a shipping financeer) addresses: Sea freight is a strictly price-sensitive industry that has used _centuries_ to find grey areas in which to shave a penny. Newcomers are chewed up and spat out. The story's protagonist, in their attempt to become a "shipping man", declines from being a wealthy fund manager to a broke divorcee.
My day job is at a freight-related SaaS. There is a lot of opportunity in freight outside of running your own asset-heavy line. We help optimise allocation. It's lucrative for us and for our customers.
But if I were to start such a thing, I'd find some "logistic managers" on LinkedIn and interview them about their pain. I expect finding a ship to send cargo would not usually be at the top of their list ...
... but it actually might be, right now, in the short term. Sea freight prices have soared over COVID, and are now 7+ times higher than before[1]. For someone determined to follow the romance of shipping, today might be a better time than most.
(I suspect you're viewing the "flex" pricing).