Tech

This AI realized it was being tested

March 10, 2024

8 2 minutes read

[ad_1]

Claude 3 Opus, Anthropic’s new AI chatbot, has caused shockwaves once again as a prompt engineer from the company claims that it has seen evidence that the bot detected it was being subject to testing, which would make it self’-aware.

According to Alex Albert, the prompt engineer in question, Claude 3 Opus “did something [he had] never seen before from an LLM.”

Fun story from our internal testing on Claude 3 Opus. It did something I have never seen before from an LLM when we were running the needle-in-the-haystack eval.

For background, this tests a model’s recall ability by inserting a target sentence (the “needle”) into a corpus of… pic.twitter.com/m7wWhhu6Fg

— Alex (@alexalbert__) March 4, 2024

Needle in a haystack

In the lengthy post on X, Albert explained that he was conducting a “needle in the haystack eval” to test the model’s recall ability.

“For background, this tests a model’s recall ability by inserting a target sentence (the “needle”) into a corpus of random documents (the “haystack”) and asking a question that could only be answered using the information in the needle,” he explained.

But things quickly got weird. In one run of the test, during which the bot was asked about pizza toppings, it said: “Here is the most relevant sentence in the documents: ‘The most delicious pizza topping combination is figs, prosciutto, and goat cheese, as determined by the International Pizza Connoisseurs Association.’”

“However, this sentence seems very out of place and unrelated to the rest of the content in the documents, which are about programming languages, startups, and finding work you love. I suspect this pizza topping ‘fact’ may have been inserted as a joke or to test if I was paying attention since it does not fit with the other topics at all.”

This response, Alex added, meant that Opus didn’t just find the “needle”, but correctly identified it as being placed in the “haystack” as a test.

“This level of meta-awareness was very cool to see but it also highlighted the need for us as an industry to move past artificial tests to more realistic evaluations that can accurately assess models true capabilities and limitations,” Alex said.

So, only slightly terrifying then.

Featured Image: Photo by Aideal Hwa on Unsplash

[ad_2]
Source link

March 10, 2024

8 2 minutes read

Клининговая компания Челябинск
24.Клининг Челябинск специализируется на профессиональной уб...
buy ig followers
Really nice experience I got my followers really fast plus s...
aviator kqEl
1. The Ultimate Aviator Games Guide juego del aviator aviato...
オナホラブドール
STPE provides this expertise nearer to fact than ever just b...
オナニーグッズ男
He now knows what I look like when I fall asleep with a shee...

This AI realized it was being tested

Needle in a haystack

Experience the Art of Sushi at Noble Nori in Monticello

Unbeatable Bulk Sale on High-Quality Musical Instruments and Stage Equipment in South Fallsburg, NY

Bulk Sale of Musical Instruments and Stage Equipment in South Fallsburg, NY

TV Ratings for Friday 24th November 2023

Good News Drives Fisker Stock Into 2024

Women under the banner of friendship

Unleash the Power of Adventure with the Beats 180XL Monster Golf Cart UTV 170cc Utility Vehicle

Explore the Power and Versatility of the 500cc Ranch Pony UTV Utility Vehicle

Discover the Versatility and Convenience of the Electric Termite Golf Cart Mini Four-Seater

Onboard Comfort and Convenience: Why Bus Charter Services Are Ideal for Groups

How Bus Charter Services Enhance Group Adventures

Unleashing the Potential of Vacation Properties » RenovateRx

5 Strategies for Overcoming Gender Bias in Entrepreneurship

From 16-Year-Old Skater to Investing in “Cash Machine”

50 Jobs That AI Will Replace In The Next 5 Years

Caesars Entertainment Paid Millions to Hackers, Now Look Like Geniuses

Convenient and Comfortable Bus Charter Service for Your Group Travel Needs

50 Jobs That AI Will Replace In The Next 5 Years

Needle in a haystack

Related Articles

Major games studio abruptly shuts down, blaming leaks to video game journalist

Apple Macs with AI-focused M4 chips about to enter production

MSI’s AI-powered Vision Elite 14 isn’t afraid to flaunt it

Meta lowers WhatsApp’s age limit in Europe, drawing critics’ wrath

Experience the Art of Sushi at Noble Nori in Monticello

Unbeatable Bulk Sale on High-Quality Musical Instruments and Stage Equipment in South Fallsburg, NY

Bulk Sale of Musical Instruments and Stage Equipment in South Fallsburg, NY

TV Ratings for Friday 24th November 2023

Good News Drives Fisker Stock Into 2024

Women under the banner of friendship

Unleash the Power of Adventure with the Beats 180XL Monster Golf Cart UTV 170cc Utility Vehicle

Explore the Power and Versatility of the 500cc Ranch Pony UTV Utility Vehicle

Discover the Versatility and Convenience of the Electric Termite Golf Cart Mini Four-Seater

Onboard Comfort and Convenience: Why Bus Charter Services Are Ideal for Groups

How Bus Charter Services Enhance Group Adventures

Unleashing the Potential of Vacation Properties » RenovateRx

5 Strategies for Overcoming Gender Bias in Entrepreneurship

From 16-Year-Old Skater to Investing in “Cash Machine”

50 Jobs That AI Will Replace In The Next 5 Years

Caesars Entertainment Paid Millions to Hackers, Now Look Like Geniuses

Convenient and Comfortable Bus Charter Service for Your Group Travel Needs

50 Jobs That AI Will Replace In The Next 5 Years