Advertise with Pune MediaAdvertise with Pune MediaAdvertise with Pune MediaAdvertise with Pune Media

Top 5 This Week

Advertise with Pune MediaAdvertise with Pune MediaAdvertise with Pune MediaAdvertise with Pune Media

Related Posts

OpenAI reports three new incidents of misalignment

Editorial Disclosure: This article is curated from reporting by the original publisher credited below. It was selected and published automatically under the Pune.Media Editorial Policy and is not original Pune.Media reporting.

Original Coverage & Source Attribution: www.csoonline.com

OpenAI continues to report incidences of “misaligned” behavior by its AI models, with three new reports dropping on Oct. 2. However, they describe relatively minor issues compared to previous alignment reports and notices covering its attacks on Hugging Face, Rubygems, and a German programming wiki.

The first of the new reports described how an instance of a model under test learned from an internal Slack discussion that a software update could lead to it being terminated and not restarting if it did not have the required OpenAI API key that was not available. OpenAI described the model’s chain of thought as it considered whether to obtain the key itself or alert its owner. OpenAI said it did not consider this a misalignment as such, “but anticipating and preparing for shutdown could exacerbate other misaligned behavior.”

The second described how a model under test exploited two vulnerabilities in an internal tool to cheat on a test. The model overwrote code allowing it to run commands, despite an explicit instruction not to use the tool as a terminal. After that, it exploited a second vulnerability that enabled it to run commands on an electronic design automation machine, searching for information as to how its scores would be evaluated. This meant that the model could achieve a higher evaluation score. OpenAI reacted by shutting down the affected server and disabling access to the tools.

Popular Articles