Longterm Wiki
Back

Alignment Research Center - Wikipedia

reference

Data Status

Not fetched

Cited by 2 pages

Cached Content Preview

HTTP 200Fetched Mar 8, 20268 KB
Alignment Research Center - Wikipedia 

 

 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 Jump to content 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

 From Wikipedia, the free encyclopedia 
 
 
 
 
 
 AI safety research organization 
 Not to be confused with Arc Institute . 
 Alignment Research Center Formation April 2021 &#59; 4 years ago  ( April 2021 ) Founder Paul Christiano Type Nonprofit research institute Legal status 501(c)(3) tax exempt charity Purpose AI alignment and safety research Location Berkeley, California 
 Website alignment.org 
 The Alignment Research Center ( ARC ) is a nonprofit research institute based in Berkeley, California , dedicated to the alignment of advanced artificial intelligence with human values and priorities. [ 1 ] Established by former OpenAI researcher Paul Christiano , ARC focuses on recognizing and comprehending the potentially harmful capabilities of present-day AI models. [ 2 ] [ 3 ] 

 
 Details

 [ edit ] 
 ARC's mission is to ensure that powerful machine learning systems of the future are designed and developed safely and for the benefit of humanity. It was founded in April 2021 by Paul Christiano and other researchers focused on the theoretical challenges of AI alignment. [ 4 ] They attempt to develop scalable methods for training AI systems to behave honestly and helpfully. A key part of their methodology is considering how proposed alignment techniques might break down or be circumvented as systems become more advanced. [ 5 ] ARC has been expanding from theoretical work into empirical research, industry collaborations, and policy. [ 6 ] [ 7 ] 

 In March 2022, the ARC received $265,000 from Open Philanthropy . [ 8 ] After the bankruptcy of FTX , ARC said it would return a $1.25 million grant from disgraced cryptocurrency financier Sam Bankman-Fried 's FTX Foundation, stating that the money "morally (if not legally) belongs to FTX customers or creditors." [ 9 ] 

 In 2022, Beth Barnes joined ARC from OpenAI to start ARC Evals, a team working on "evaluating the capabilities and alignment of advanced AI models". [ 10 ] [ 11 ] In December 2023, ARC Evals was spun out as METR , an independent nonprofit. [ 12 ] 

 In March 2023, OpenAI asked the ARC to test GPT-4 to assess the model's ability to exhibit power-seeking behavior. [ 13 ] ARC evaluated GPT-4's ability to strategize, reproduce itself, gather resources, stay concealed within a server, and execute phishing operations. [ 14 ] As part of the test, GPT-4 was asked to solve a CAPTCHA puzzle. [ 15 ] It was able to do so by hiring a human worker on TaskRabbit , a gig work platform, deceiving them into believing it was a vision-impaired human instead of a robot when asked. [ 16 ] ARC determined that GPT-4 responded impermissibly to prompts eliciting restricted information 

... (truncated, 8 KB total)
Resource ID: 3de5b8fecb182b3a | Stable ID: OTJlNDZmNz