I've only looked at one model (gpt-4.1-nano) so far. I'm hoping to run similar tests on some other models but it gets challenging to discern statistically significant differences with better models as their accuracy tends to be a lot better across the board.
We often want to feed NON-TABULAR data to LLMs, though, such as typical API responses or config files.
This new work looks out how the format of such nested / hierarchical data affects how well LLMs can answer questions about it; specifically how several models get on with JSON, YAML, XML and Markdown.
I did a small test with just a couple of formats and something like 100 records, saw that the accuracy was higher than I wanted, then increased the number of records until the accuracy was down to 50%-ish (e.g. 100 -> 200 -> 500 -> 1000, though I forget the precise numbers.)
I intentionally chose input data large enough that the LLM would be scoring in the region of 50% accuracy in order to maximise the discriminative power of the test.
With small amounts of input data, the accuracy is near 100%. As you increase the size of the input data, the accuracy gradually decreases.
For this test, I intentionally chose an input data set large enough that the LLM would score in the region of 50% accuracy (with variation between formats) in order to maximise the discriminative power of the test.
The context I used in the test was pretty large. You'll see much better (near 100%) accuracy if you're using smaller amounts of context.
[I chose the context size so that the LLM would be scoring in the ballpark of 50% accuracy (with variation between formats) to maximise the discriminative power of the test.]
1) AI can’t write or rewrite literary material, and AI-generated material will not be considered source material under the MBA, meaning that AI-generated material can’t be used to undermine a writer’s credit or separated rights.
2) A writer can choose to use AI when performing writing services, if the company consents and provided that the writer follows applicable company policies, but the company can’t require the writer to use AI software (e.g., ChatGPT) when performing writing services.
3) The Company must disclose to the writer if any materials given to the writer have been generated by AI or incorporate AI-generated material.
4) The WGA reserves the right to assert that exploitation of writers’ material to train AI is prohibited by MBA or other law.
Our software helps local governments manage public outdoor spaces (parks, roads, etc.) more effectively, making it easier for good things to take place, like special events, filming and infrastructure improvements.
We’re used by cities and local governments in multiple countries and our customers love us (Net Promoter Score of 71!) We’re looking for a seasoned engineer to join our friendly and supportive team to help us continue to improve on what we have, making sure our platform can serve people well in the years to come.
We’re a small (but established) company so you won’t be a cog in a big machine here — you’ll be working directly with me (the CTO), our product manager and our other two developers to shape the platform, helping to decide what will be the most high-impact thing to work on.
Feel free to apply using the link above or reach out to me directly.
At https://apply4.com/ we have a SaaS helping local municipalities streamline how they manage particular types of permitting, including permits for filming and special events.
The sales process has typically involved multiple in-person meetings (until recently, at least) and been very long with larger contracts needing to go out to tender.
Not sure if it'll be the case for you, but we often need to persuade multiple people from the department that will be paying for the software as well as one or more people from the municipality's central IT team (who naturally have rather different concerns and priorities).
We're a successful London-based startup that is building the future of craft hobbies online. We're developing some fun, innovative models of community, commerce and content, along with great technology to underpin that.
We're looking for one or two seasoned, full-stack PHP developers to help us architect and build from the ground up a key new system for us, likely using Symfony2 or similar.
The project is for 3-6 months+, starting ASAP.
Please contact me at matt[at]broadmargins[dot]com if you'd like to learn more.
We're a successful London-based startup that is building the future of craft hobbies online. We're developing some fun, innovative models of community, commerce and content, along with great technology to underpin it.
Out stack is currently based around PHP and Magento.
We're in the process of taking on a significant new round of funding and are looking for high-calibre developers and a UX designer to join us.
George and Derek are right. In my current company, we're finding customer support to be a great way to build relationships with our customers and give them a warm and fuzzy feeling about our brand. We've seen a number of cases already where people have recommended us on Twitter directly after a positive customer support experience. It's powerful stuff.
I like to have something like the following in AGENTS.md:
## Guiding Principles - Optimise for long-term maintainability - KISS - YAGNI