This is so unfortunate but a clear illustration of something I've been thinking about a lot when it comes to LLMs and AI. It seems like we're forgetting that we are just handing our data over to these companies on a solver platter in the form of our prompts.
Disclosure that I do work for Tonic.ai and we are working on a way to automatically redact any information you send to an LLM - https://www.tonic.ai/solar
I think that there is a science here that people are working to discover. Prompt engineering is slowly turning into its own language and the more it's studied like this I think the more validity there will be to calling it engineering.
I interviewed Bing about how it uses synthetic data in its training.
TLDR; If asked a question on a topic that it feels there’s a lack of quality data to train on, Bing builds its own GANs to generate more samples for training in order to provide higher quality outputs.
I really like this point, that fast thinking is shallow. I think there should be more respect for having to think deeply about things and process before coming to conclusions. I think that is more of a sign of intelligence than just responding to something right away.
This would be the ideal - the advantage that this has over google is that you don't have to wade through so many hits to find the page that has your answer and then read through the whole page to find your answer.
Having the concise paragraph answer of chatGPT with sources to cross reference is the most ideal.
This is an interesting idea, however, I find it difficult to find time to even sit and scroll through Twitter and read news-based newsletters. It would be nice to have long form updates from friends and write my own as well, but I think with the pace of living this is not possible - which is unfortunate!
Very good point about the implications of large language models for the actual act of dating! As if there weren't already enough reasons to be suspicious of people ("people") you are interacting with on those platforms...
You also make a good point about replicating real data - I think this is another area where more tools need to be developed to safe guard against consequences. In addition to the obvious challenges to plagiarism softwares LLMs pose, there are definitely opportunities here to develop privacy detecting softwares should an org want to use GPT-3 or ChatGPT for synthetic data.
Recently I was reading a book by Peter Diamandis and Steven Kotler about converging technologies that referenced the Luddites' reaction to the invention of the loom, fearing it would lead to the destruction of jobs, when ultimately the innovation lead to more economic opportunity. I think we'll find that the fears people are having in this regard about LLMs are similar with opportunities like these to develop more digital infrastructure around their application.
ChatGPT has taken the tech world by storm, but its older cousin GPT-3 is still relevant. Being able to connect to the text completion API through python allows you to use the large language model to generate synthetic data with bespoke distributions.
The application is limited, however, as the lack of on-prem deployment limits your ability to show the model your proprietary data to learn from.
"The identity ecosystem" he describes is most exciting to me. There is so much that can be attached to our on-chain identity beyond just our net worth and jpegs. This year the first house was sold on-chain, and I believe that's just the beginning of these immutable irl transactions.
I"m curious if artists can opt out of their songs being included in those that have these capabilities. This seems like another way that a streaming service is exploiting artists to me. It's cool technology definitely, but there are some problematic implications for sure.
I don't know, I feel like as a programmer these technologies make a lot of sense. AIs like this have been being developed for so many decades it's not at all surprising that we are finally at a place where they feel like we're talking to another human. Though I have to admit it's still kind of scary, just not unbelievable.
It’s so difficult to build an unbiased model to classify a rare event since machine learning algorithms will learn to classify the majority class so much better. This blog post shows how a new AI-powered data synthesizer tool, Djinn, can upsample synthetic data even better than SMOTE and SMOTE-NC. Using neural network generative models, it has a powerful ability to learn and mimic real data super quickly and integrates seamlessly with Jupyter Notebook.
Full disclosure: I recently joined Tonic.ai as their first Data Science Evangelist, but I also can say that I genuinely think this product is amazing and a game-changer for data scientists.
Happy to connect and chat all things data synthesis!
Wow, losing your job is already painful enough, but I can't even imagine the trauma of also facing losing the life you've built in this country on top of that. 90% of immigrant layoffs being H-1B holders is a staggering statistic.
Wow, this is an effect of our financial crisis that I hadn't considered before... I'd be very interested to see the data on the proportion of people who have been laid off who were sponsored for H-1B visas.
These types of things are so frustrating as they go against the core ethos of cryptocurrency. This announcement from Greyscale has tones of big banks and it's scary to see the crypto industry to trend this way so quickly.
On the other hand, security concerns should be taken seriously as well, especially when it comes to potential runs... interested to hear others' thoughts on this.
It's interesting because I don't know if it's possible for anything to truly replace Twitter. I think what will happen is people will splinter off into online forums where they feel most comfortable based on the level of moderation on each site and who is there. Twitter was unique in that everyone was on it, sadly I don't think we'll ever have that again.