LWiAI Podcast #208 - Claude Integrations, ChatGPT Sycophancy, Leaderboard Cheats

Podcast

LWiAI Podcast #208 - Claude Integrations, ChatGPT Sycophancy, Leaderboard Cheats

0:00

-1:55:25

LWiAI Podcast #208 - Claude Integrations, ChatGPT Sycophancy, Leaderboard Cheats

Anthropic lets users connect more apps to Claude, OpenAI undoes its glaze-heavy ChatGPT update, and more!

Last Week in AI

May 10, 2025

Our 208th episode with a summary and discussion of last week's big AI news!
Recorded on 05/02/2025

Hosted by Andrey Kurenkov and Jeremie Harris.
Feel free to email us your questions and feedback at contact@lastweekinai.com and/or hello@gladstone.ai

Join our Discord here! https://discord.gg/nTyezGSKwP

In this episode:

OpenAI showcases new integration capabilities in their API, enhancing the performance of LLMs and image generators with updated functionalities and improved user interfaces.
Analysis of OpenAI's preparedness framework reveals updates focusing on biological and chemical risks, cybersecurity, and AI self-improvement, while tone down the emphasis on persuasion capabilities.
Anthropic's research highlights potential security vulnerabilities in AI models, demonstrating various malicious use cases such as influence operations and hacking tool creation.
A detailed examination of AI competition between the US and China reveals China's impending capability to match the US in AI advancement this year, emphasizing the impact of export controls and the importance of geopolitical strategy.

Timestamps + Links:

Tools & Apps

(00:02:57) Anthropic lets users connect more apps to Claude
(00:08:20) OpenAI undoes its glaze-heavy ChatGPT update

(00:15:16) Baidu ERNIE X1 and 4.5 Turbo boast high performance at low cost
(00:19:44) Adobe adds more image generators to its growing AI family
(00:24:35) OpenAI makes its upgraded image generator available to developers
(00:27:01) xAI’s Grok chatbot can now ‘see’ the world around it

Applications & Business:

Projects & Open Source:

(00:47:14) Alibaba unveils Qwen 3, a family of ‘hybrid’ AI reasoning models
(00:54:14) Intellect-2
(01:02:07) BitNet b1.58 2B4T Technical Report
(01:05:33) Meta AI Introduces Perception Encoder: A Large-Scale Vision Encoder that Excels Across Several Vision Tasks for Images and Video

Research & Advancements:

(01:06:42) The Leaderboard Illusion
(01:12:08) Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
(01:18:38) Reinforcement Learning for Reasoning in Large Language Models with One Training Example
(01:24:40) Sleep-time Compute: Beyond Inference Scaling at Test-time

Policy & Safety:

(01:28:23) Every AI Datacenter Is Vulnerable to Chinese Espionage, Report Says
(01:32:27) OpenAI preparedness framework update
(01:38:31) Detecting and Countering Malicious Uses of Claude: March 2025
(01:46:33) Chinese AI Will Match America's

Discussion about this episode

Ready for more?

#nojs-banner { position: fixed; bottom: 0; left: 0; padding: 16px 16px 16px 32px; width: 100%; box-sizing: border-box; background: red; color: white; font-family: -apple-system, "Segoe UI", Roboto, Helvetica, Arial, sans-serif, "Apple Color Emoji", "Segoe UI Emoji", "Segoe UI Symbol"; font-size: 13px; line-height: 13px; } #nojs-banner a { color: inherit; text-decoration: underline; } This site requires JavaScript to run correctly. Please turn on JavaScript or unblock scripts