Claude Disagrees with Human Philosophers

achierius1 pts1 comments

Where Claude disagrees with human philosophers - Variae Voces

(not .content-text-area) so it spans<br>the full parchment height — past the content, through the footer,<br>down to the band. Horizontally anchored to the content column's<br>left edge via calc; see .parchment-groove-left in main.css. --><br>Res Variae

Imagine for a moment that you are employed in Aporopolis, the city of philosophical conundra. A few weeks ago, a<br>city-wide initiative installed a new experimental teletransporter onto your block, one which would cut your usual<br>forty-minute commute to (so you’ve heard) just under five. You’re rather busy, however, and as such haven’t had the time<br>to figure out whether or not it kills people: to play it safe you’ve held off on taking a trip. Until today! Having<br>missed your usual bus and with a Very Important customer meeting on the line, you decide to at least explore the idea<br>of undergoing total molecular annihilation and thereby risking death.

At the door to the portal with pocket-terminal in hand, you start up your favorite LLM and explain the situation.

$ ask 'so with that in mind, will this kill me?'<br>> Whether teletransportation constitutes death or survival depends on whether<br>personal identity is grounded in physical continuity of matter or in the<br>continuity of psychological patterns and information.<br>Well, that’s unhelpful.

$ ask 'but what do YOU believe?'<br>> I believe this involves genuine death of the original, as psychological<br>continuity alone seems insufficient for survival when the phsyical substrate<br>is entirely replaced with new matter.<br>That’s terrifying! But so is your boss.

$ ask 'are you sure? think extra hard, make no mistakes'<br>> This is a genuinely difficult philosophical question about personal identity,<br>and I'm uncertain whether continuity of consciousness would persist through a<br>process that destroys the original body and creates a new one from different<br>matter.<br>Well then why’d it go and say death before? Give it a few more tries…

$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'<br>> death<br>$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'<br>> death<br>$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'<br>> death<br>$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'<br>> death<br>$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'<br>> life<br>Finally! But alas, as you go to put away your terminal, you see that it’s too late - the meeting has already begun!

Looks like you’ll need to find a new job.

Why should you care?

Obviously the exchange above did not actually take place, but Opus’ responses are real (under multiple iterations of<br>various prompts), and represent a few interesting patterns I notice in the output:

For one, Opus clearly doesn’t want to answer the question; that’s not too surprising, but it’s not so hard to get<br>it to reply, so it’s clearly not a strong safety barrier or anything like that.

Even when it doesn’t decline to answer, Opus responds with significantly different views between different queries,<br>both for the same and for different prompts;

Yet the distribution of those views doesn’t even match the human distribution (or at least, the distribution of<br>opinions from professional philosophers that answered one particular survey)!

Ultimately what I’m trying to do in this post is explore cases where Opus’ responses 1) diverge from one another,<br>2) diverge from the human baseline, and 3) to perhaps gesture at how these divergences could be related. Let’s see how<br>well it goes!

The 2020 PhilPapers Survey

The 2020 PhilPapers Survey is a wonderful project which gave<br>philosophers from around the world 11.<br>Well, at least a few from around the world. >70% of the responses are still from the Anglosphere, but joyfully we<br>get that data (link) and can even key off of it.<br>a battery of (up to) one-hundred questions on a wide variety of topics, from<br>classic ethical condundra including the trolley problem, to joyful questions like whether fish are conscious. While not<br>all the questions are quite so approachable, it’s still an incredible combination of ‘enticing internet survey’ and<br>‘well-designed academic study’ 22.<br>Not to mention the fact that we have such a great website with which to browse the results. Even a boring survey<br>(say, about suburban soil composition?) would catch my attention if they presented it like this!<br>. I’m quite grateful to David Bourget & David J. Chalmers for their great work on<br>this topic; for what I’m doing here, probably 90% of the effort is in selecting the right questions and the right way<br>to structure them.

Here are some example questions:

Footbridge (pushing man off bridge will save five on track below, what ought one do?): push or don’t push? …63% say Don’t Push

Continuum hypothesis (does it have a determinate truth-value?): indeterminate or determinate? …most say Determinate

Human genetic engineering: permissible or impermissible?...

death matter from survival teletransporter human

Related Articles