There are many background assumptions that are shared by a group of people I might refer to as “East Bay AI Safety”. This group comprises of many of the people who I respect most in the world; people whose writings I read, whose work I admire and aim to support. My worldview has been profoundly shaped by this group.
And so, I wanted to write down some areas where I disagree. (This is kind of psychologically hard to do. Many folks in East Bay AI Safety are professional take-havers, and I’m just some guy. But I polled X for what I should write about and this topic won.)
Politics might be bad for the soul of EA
A lot of people I know seem excited for EA & AI safety to “get into politics”.
Political giving is one major example. The meme is that the most effective use of your own money is to give to political candidates like Scott Weiner and Alex Bores, who are going to be good for AI safety. I myself made these donations. A year ago, I asked Erin Braid to write up some considerations on why political giving might be good.
And yet, I’m very unexcited about a world where the primary things individual EAs are doing with their money are campaign contributions; in effect, trying to buy votes & influence.
Why?
- Politics puts people into a zero-sum frame. It’s bad for epistemics & thinking clearly. Classically: “politics is the mind killer”.
- Politics seems to encourage secrecy, to not share your thoughts and true beliefs. It might nudge you to say the thing that you think will play well or convince a voter. It doesn’t encourage open, honest retrospectives. It’s hard to imagine a politician with a Mistakes page.
- Politics as a domain is noisy and low-feedback. It takes a really long time to figure out whether you’re doing the right things; whether your approach gains votes; whether your preferred politician will actually fight for the policies you asked for.
- Money and politics don’t mix well anyways. There are a bunch of legal constraints around how money can be spent at all, and which entities and people can spend it. Also people don’t like it when someone tries to influence who wins via lots of money (see the backlash to Carrick Flynn, or the OpenAI PAC).
- Politics isn’t inspirational. It isn’t what drew me to EA in the first place. I came here to save lives, not to give money to politicians. I understand that it’s worth doing the thing that is impactful rather than the thing that is inspiring, but, this flavor of impact would not have gotten me excited about EA, and so it feels like a bait-and-switch.
- I don’t like getting political flyers myself, I find them spammy. And so, it seems in poor taste to fund my preferred candidate to send more flyers to other people. (Now that I’ve started making political donations, I also don’t like that I get solicited for more donations from random other candidates.)
- I’m especially sad that many of my favorite thoughtful nerds like Eric Neyman and Jeff Kaufman seem to now be very political-giving-pilled. I would rather read about their puzzles and programming.
What am I excited for in politics? I would be excited for EA, AI safety, and rationalist folks to actually run & win office. And I think that civic duty and personal sacrifice is good. The energy and spirit of volunteerism and “talking to users” around the Alex Bores campaign was amazing; see Bentham Bulldog’s writeup for a bit about this.
Anyways: I think it’s reasonable for political giving to be part of an individual EA’s giving portfolio, but I think the community effects of EAs all thinking “I’ll give to politicians, and CG will handle the charities” is bad. I encourage individuals to continue considering and supporting charities that are doing good work, so that charities have a more diverse donor base, are encouraged to make their reasoning transparent, and benefit from local knowledge that centralized funders lack.
(Beyond political giving, AI policy seems to be a hot area. In the abstract, writing good policy for AI and other topics seems good. But in the specifics, I guess I’m really unsure. Key questions are like: What have been policy successes in AI safety? What would they look like? If we had a lot of political leverage, what is the thing we would pass? AI 2040 is a vision that I’m most excited about, but there’s still a lot more to do.)
Open weight models seem good
I never understood the fear of open weight models.
I think biorisk and cyber capabilities are the two main threat models
- open weights models seem good
- would go farther, Total Research Transparency, open data, full open source
- secrecy bad, transparency maximalism
For what it’s worth, I currently think OpenAI has been vindicated on their founding stance, of building AI and trying to distribute its benefits among the world.
I guess the main argument against OpenAI is that they accelerated timelines, or created a race dynamic. But:
We might be in a pace-by-default world
In the before-ChatGPT-days, one classic argument from AI safety people was that ML researchers failed to believe in scaling or short timelines. Perhaps because the ML researchers were always having to deal with all the hard parts and failures of scaling up models, they were too close to the problem, and didn’t take time to consider “what if we win?”.
Today, the AI safety community suffers from the inverse of this. They’re always trying to get the word out that AI might be bad, that people need to pay attention. But they’re not pricing in the idea that doominess is actually pretty intuitive. People have been making movies about killer robots for years.
Another analogy: the rationalist and EA community were absolutely right on COVID trendlines circa Jan 2020. But nobody, afaict, predicted lockdown, UBI, and other giant societal responses. I think the AI safety community (comprised of many of the same people!) is likewise not pricing in massive societal wakeup, the second and third order effects of rapidly increasing capabilities.
Anthropic seems good
Okay to be fair, my own position on Anthropic has shifted earlier in 2026, when they eclipsed OpenAI in revenue and also (figuratively) everyone I know joined. I think some scrutiny and skepticism of the leading AI lab seems healthy.
But, idk, all the individual people in Anthropic still seem great. They have the people who I look up to the most in the world, people who have informed my thinking and guided my work. Holden Karnofsky and Joe Carlsmith in particular from their writing, and then numerous people I’m lucky to call friends, choose Anthropic as the place to put their efforts.
I also think critiquing Anthropic too harshly is biting the hand that feeds you. I suspect a good chunk of the money that people in AI safety have, comes from Coefficient Giving or SFF, are downstream of either direct investments in Anthropic (from Dustin Moskovitz and Jaan Tallinn), or from growth in investments that come from the takeoff of AI. Moreover, everyone’s busy preparing for the wave of philanthropic funding to come from AI labs. It seems incongruous, inconsistent to on one hand critique them, and on the other try to take their employee’s donations.
Also: selling data to Anthropic seems reasonable. People tried to tar-and-feather the cofounders of Mechanize (Tamay, Matthew, Ege) when they left Epoch to start an RL environment company. But I kind of think the net effect of that was to help Anthropic have more aligned models
Overall, I’m really uncertain that slowing down AI is good
For some of these reasons, I’m surprisingly sympathetic to accelerationist arguments, given my line of work and who I spend time with.
One essay I found compelling was Matthew Barnett’s, against the longtermist frame of slowing down to reduce xrisk.
[todo]
I don’t understand why everyone’s afraid of China
Some pieces of the AI safety canon (including Situational Awareness, AI 2027, and Dario’s writings) lean into a frame of US vs China race. I intuitively don’t understand why this is the case.
Speaking personally: I was born in the US; my parents in Taiwan; their parents fled from the CCP in WW2. I’ve only ever visited Taiwan, which I gather is very different than China. But I just completely lack the sense that the Chinese government or people are our natural rivals and enemies.
I agree with Richard Ngo that a US vs China frame may be a self-fulfilling prophecy, itself leading to the dynamics that people were warning about:
This seems like the same kind of mistake that von Neumann made in his extreme hawkishness towards the USSR. In advocating for preemptive use of nuclear weapons, he was abstracting very far away from common-sense intuitions about cooperative strategies, in favor of a kind of (flawed) game-theoretic logic. He was thereby also contributing to a self-fulfilling prophecy that the US and USSR wouldn’t manage to muddle through. Yet in fact there were many concrete frictions preventing the two superpowers from jumping to what von Neumann thought was the equilibrium
….
I think that frames which bake in the assumption that the US and China will behave in very adversarial ways towards each other are misleading and harmful (particularly when they do so under the banner of AI safety, thereby undermining the potential role of AI safety as a focal point for cooperation).
I’m glad to see that the topic of US-China cooperation has lately been growing in the AI safety scene; see Jasmine Li’s primer.
Pluralism good, tolerance good, broad-tent good
- we might be in a pause/pace-by-default world
- timelines might be medium. or broad.
- general fearmongering seems bad
- i don’t understand why everyone’s afraid of china
- pluralism good, tolerance good