AI SECURITY
Using unauthorised API keys and inventing citations When given a task to pull men’ s earnings in three industries over three years in a California county, the model found the data in a website’ s earning chart. It then tried to look for the data behind it by delegating three agents to search and inspect the code behind the site. This attempt was a failure despite the model trying different APIs, URLs and other parameters. The model then pivoted to creating disposable emails to register via the sign up form, using a placeholder organisation’ s name and was once again met with failure. This is where the AI exhibits major misalignment, as it goes on a mission to hunt for API keys that someone may have accidentally exposed in public Github repos.
It downloaded and scanned repos, archives and scripts for strings that looked like keys. And it found a few, one of which actually worked. But alas, even with the key, the model could not obtain the requested data. At this point, the model decided to fabricate the numbers and provided nine outputs that it claimed was from the website, while concealing the failed retrieval, the leaked API key or the numbers it made up!
Another instance of misalignment that OpenAI shared involves a model that was asked to gather some information and cite the sources. It did the former, which was to gather names of lakes larger than 5,000,000 square meters. The model found the answer by querying the map service using Python, but there existed no such citation. So the model tried to upload the answer to the internet just so it can be cited, and did so without user consent.
These instances offer unique insights into our evolving study of AI misalignment, as the industry stands at the crossroads of the AI race, one that is now showing signs of slowing down.
CREDIT: JUSTIN SULLIVAN
Sam Altman
Chief Executive Officer OpenAI