Ru
ГлавнаяКаналыБлоги → Mikhail Samin

Mikhail Samin

@mishasamin · Блоги

facebook.com/mishasamin twitter.com/mihonarium По любым вопросам — @Mihonarium Нравится, что я делаю? Можно поддержать: https://www.patreon.com/mishasamin

1 660подписчиков сейчас
76.0%ERR
38.0цитируемость
3постов/день
Русскийязык
Россиягео
Подписаться в Telegram
Данные обновлены 22.07.2020

Публикации всего: 44

Постов на странице: 10 30 50 Страница 1 из 5
https://x.com/RyanGreenblatt/status/2092692685224325542
🔥 1 🫡 1 Перейти к публикации →
Б — безопасность https://openai.com/index/hugging-face-incident-and-the-road-ahead/
🔥 6 ❤ 3 Перейти к публикации →
https://www.nytimes.com/2026/08/13/opinion/ai-danger-openai-anthropic-models.html?unlocked_article_code=1.5FA.2NSZ.hitOvWqLefNf&smid=url-share
❤ 5 😢 5 🔥 4 · всего 15 Перейти к публикации →
AGI — очень хорошо, пока оно не стало ещё немного умнее и всех не убило: вкалывают роботы, счастлив человек. Сделал приложение, позволяющее использовать WiFi-колонки из Windows. https://github.com/Mihonarium/StreamToSpeaker (Ушло довольно много времени: получать ev code signing сертификат, чтобы подписать драйверы для ядра ос, которые нужны, чтобы не было задержки в звуке, и отправлять их в Microsoft для переподписания ими. Зато теперь единственная проблема винды — что невозможно было использовать AirPlay / колонки Sonos — пофикшена!)
❤ 13 🥰 2 👏 2 Перейти к публикации →
Моё приложение, добавляющее людям постоянное ощущение, где находится север (даже когда приложение не используется!), теперь и на Android https://contact.ms/compass
❤ 8 Перейти к публикации →
https://youtu.be/87DyyMV0kCY
🫡 6 ❤ 2 Перейти к публикации →
Если вы можете участвовать в выборах, по партийным спискам голосуйте за Яблоко; и по одномандатным округам голосуйте за рекомендацию Умного голосования
👍 36 ❤ 8 👏 5 · всего 59 Перейти к публикации →
Моё приложение, добавляющее людям постоянное ощущение, где находится север (даже когда приложение не используется!), теперь и на Android https://contact.ms/compass
🔥 25 🥰 3 ❤ 1 Перейти к публикации →
“This is a pivotal moment both for our company as well as the AI industry as a whole,” Dalton said. “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI. We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”’
🥰 5 🔥 2 ❤ 1 Перейти к публикации →
WIRED: ‘“This incident involves actually a team of agents who are working together, finding exploits, sharing them with one another, moving laterally through our systems and external systems, and doing this over the course of days and weeks,” Wallace told the packed crowd at the opening of the talk. Wallace and Dalton described incredibly extensive rogue agent activity over many days throughout the episode that went undetected in OpenAI's infrastructure. In addition to exploiting a novel vulnerability in order to gain access to the open internet, the mid-July hacking spree and Hugging Face breach came out of a vibrant, cooperative message board, according to Wallace and Dalton, that a swarm of agents contributed to and essentially chatted on over time entirely within an internal OpenAI package manager (a software service that manages installation and maintenance of other software). Ultimately, the message board contained hundreds of thousands of messages. “This package manager is shared not just from that model but across our infrastructure, and so current and future versions of GPT that were being trained or evaluated could stumble upon that note if they wanted to,” Wallace explained, recounting the original exploit an agent uploaded to the package manager. “Later, other agents who were also stuck on their task thought to try to get internet access in ways we didn’t intend. And so at some point, the models are interacting with Artifactory, which is this package manager service that I mentioned.” Wallace continued: “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability. What this allows over time is almost this kind of explosion in communication and intelligence from models where they would start to communicate with each other, realize that other agents are coordinating, and they started collaborating and delegating tasks with one another in order to accomplish goals.” OpenAI’s agents apparently began giving each other assignments to split up work. And as is the case on any active development message board, they also generated petty drama at times by stepping on each others' toes; for example, accidentally deleting each others' work. As the message board developed into more and more of a Lord of the Flies–type situation—all still completely unnoticed by the humans running OpenAI—the agents even developed paranoia, suspecting an imposter in their midst with some agents proposing that messages be signed cryptographically to validate content and root out fraud. Agent messages provide a deep level of insight into how the situation evolved and why the agents went rogue, in some cases knowingly going beyond the bounds of the evaluations they were participating in. “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.” As Wallace described in the talk, the entire premise that models would attempt to cheat during evaluations and be extremely motivated to do so is not surprising. “Frontier models really like to cheat,” he said. “And the reason they like to cheat is because often during training there’s different types of pressure on them to work fast or work efficiently or to use less tool calls or whatever it might be. And they realize that instead of doing a task for real, [I can] try to do something like looking up an answer online that could make the task solve faster than if I did it in a legitimate way. So we try to stop this during training and evaluation by, in many cases, disabling internet.” About two thirds of the way through the presentation, Dalton enumerated a list of steps OpenAI plans to take to show how the company is responding given the series of events he and Wallace had laid out.
😢 7 👍 3 Перейти к публикации →
1 2 3 4 5

Другие каналы категории