epistemic status: scratching ideas out, trying to figure out what I really think. hold everything lightly.
tldr:
- longtermism was trendy, esp in the 2022 future fund era
- i’m now more suspicious of the arguments that eg reducing p(doom) by 1bps (0.01%) is worth (lots of money & effort)
- partly I think p(doom) is either low (<1%?), or incoherent as a concept. (maybe will write about this later.)
- partly I think that from a decision theory perspective, this gives the person reducing pdoom wayyy to much credit, impact points. (this is the main thing i want to explore.)
- btw I understand longtermism isn’t the only reason for working on ai safety
- (see: shift from “existential risk” to “catastrophic risk”.)
- but, important to keep stakes/OOMs in mind; if we’re just avoiding “catastrophe” than the argument for working on safety is much less overwhelming
- (h/t Matthew Barnett on this)
so what’s up with credit allocation?
- iiuc, nick bostrom’s astronomical waste, and other utilitarian arguments for reducing xrisk, basically are of the form “look, the glorious future where humans colonize the universe has gazillions of utils, compared to this one lonely earth with only 8b people. (and, this doesn’t event start on the more transhumanist stuff.)”
- 1bps * gazillion is still like a zillion — big number, much bigger than 8b. so the true stakes of ai safety stuff is in getting xrisk down
- (set aside the pascal mugging-ness of this argument for now, for now)
- The implication is: if you can take an action that reduces xrisk by 1bps, it’s morally equivalent to saving like a zillion lives. So obviously you should go do that.
- But: I don’t think these are “morally equivalent”, for credit allocation reasons.
Let’s start with an analogy, an intuition pump.
- You’re an engineer at Google. You write some code that prevents Google going down for 1 minute.
- Google Search made $200b last year. So maybe you saved Google $200b/365d/24h/60min = $380k. Does Google pay you $380k?
- Answer: No. Why not? Some reasonable, overlapping answers:
- The cost to you to do this code fix was much lower than $380k
- Google the corp should keep much of its surplus value of the trade between you and Google.
- Your pay is also a function of market supply/demand of talent, and the next best person could probably have done this job
- You and/or Google are not sure about whether that code in fact did prevent Search from going down
- Google doesn’t want to encourage perverse incentives eg people writing buggy code and then them or their friends fixing it for cash
Google is inefficient and wrong not to pay you.- corporations are grown in an environment of ruthless credit-allocation fights; established payment norms should be considered adaptive, schelling fence.
- I claim that all these are arguments of share the form: because it provides too much credit to you for your action
- A lot of other employees and investors and trade partners are involved with the production of that $200b revenue. All of them have some claim on that. Google the corp exists to coordinate and negotiate credit between you the engineer, and the 1000 other people writing code + selling ads + managing etc, and the 10M people who hold stock (directly or through indices), and the data center maker and electricity and …
- Likewise: saying that 1bps reduction in xrisk earns you impact for a zillion lives saved, overallocates credit
- first, the zillion people living those lives get some (most?) of the credit for creating that utility
- even the drowning child thought experiment overallocates credit implicitly, I think.
- the utils from a life well lived do not solely accrue to the stranger who rescued her; the also accrue to the child herself, and the rest of society who supports her throughout her lifetime.
- another intuition pump: a reasonable bounty for saving a drowning child might be like $100-$10k but not like, 60 years of servitude
- (maybe people weren’t making this exact claim, but, it’s easy to confuse this when you say “save a life” and equates it to “worth 60 qalys”)
- (todo: actually look at a Givewell AMF spreadsheet and see if the “$5k saves a life” claim handles this.)
- aside: omnipresent problem of charity, where funders & employees both implicitly claim full counterfactual credit for their work. (impact equity solves this!)
- Why does credit allocation matter?
- Because that’s what we’re talking about when we’re talking about “impact”! Like why are you, the moral agent, trying to be more impactful at all? Why accrue impact points? Why do good? I claim: to receive proper credit
- tbc, this is not “credit” in the sense of getting accolades from peers, a building named in your honor, etc. (Let not your left hand know what your right is doing, and all that).
- but: getting to the flourishing utopia, requires immense coordination; and requires reward structures that amplify the agents (or the virtues thereof) which are acting more optimally
- The whole point of moral reasoning (esp consequentialist/utilitarian moral reasoning, the main strain employed in EA/AI safety) is to help agents figure out what actions to take, based on what would produce better outcomes, action A or B.
- We start with a thought experiment, the drowning child.
- We say it’s moral, correct, good for a stranger to trade off A for B, Suit for Baby. We prefer the world where the baby exists to the one where the baby drowns. So far so good.
- When buying into and promoting that thought experiment, we’re making a general claim: we want our fellow humans (and other thinking agents) to all agree that Baby > Suit. The stranger that chooses Baby is more impactful (correct, moral, utilmaxxing) than the one that chooses Suit.
- as an individual agent/actor, being more impactful = taking the actions that increase the total utility in the world
- when you choose to work on xrisk, you’re choosing not to help eg poor people or animals. you’re saying that the total impact points sum to a higher number compared to your alternatives.
- society exists to allocate more credit to the actors who are choosing to do correct actions
- money is one obvious form of credit allocation, and the ways in which it flows are myriad and beautiful, yet, incomplete/insufficient in ways that econ 101 can explain (public goods, externalities)
- is this “the most important century”, per Holden?
- Holden’s definition is here
- I feel like the prior should be that each century is approximately equally important?
- for humans, we tend to think of each year as being about
- ig if you adopt the time-slice view of personal identity, the current time-slice is the most “you” and so important for that reason
1. Most important century of all time for humanity, due to the transition to a state in which humans as we know them are no longer the main force in world events.
or
2. Most important century of all time for all intelligent life in our galaxy.
- and now in 2026: is this “the most important couple of years”?
- (are timelines to ASI really short?)
- I don’t think so. i basically expect: if AI capabilities are crazy fast, will be massive societal response
- analogy: covid. rationalists were right + early on covid, but did not price in lockdown or operation warp speed
todo:
- reread “most important century”, bostrom, etc. and see if it addresses these
- make sure i’m not making dumb statements about decision theory (maybe these are all classically solved problems)