In this post, I try to briefly make the case for making models good at moral philosophy. To be up front, aside from some niche areas of moral philosophy, I don’t actually think this is the most important thing to work on for most people. But, depending on the exact scope, I do think it’s near competitive. I also like it as an archetype for a broader set of interventions.
After a brief preface (related to the archetype point), the main text of the post is structured as follows:
- I say what I mean by moral philosophy. I point out some directions in particular that seem relatively important and that are arguably within the scope of moral philosophy.
- I offer a brief explanation of what I mean by making AI systems “better” at moral philosophy. (I give a more detailed account in the appendix.)
- I give some examples of the kinds of interventions I have in mind.
- Finally, I make the actual case, arguing for importance and tractability (where tractability prices in neglectedness).
I include a lengthy “appendix” with the following content:
- A collection of articles by other people making similar arguments
- Responses to some common concerns about how working on this might have bad effects on the world (like shaping model motivations in problematic ways)
- Further detail on various components of the main text
I thank Alex Kastner, Chi Nguyen, and Emery Cooper for comments on earlier drafts of this post.
Preface: What about conceptual reasoning?
As readers of this blog might know, I’m interested in making AI better at various conceptual domains. In my mind, moral philosophy is a central example of a conceptual domain. It’s also a straightforward example of a domain on which I’d like AI systems to be better. There are lots of other conceptual domains on which I’d like AI to be better, e.g., foundational decision theory questions, most conceptual aspects of AI safety research, various foundational aspects of game theory (bargaining solutions, equilibrium selection, safe Pareto improvements).
I think moral philosophy, in its broadest form, is probably not the most leveraged conceptual domain to work on making AIs better at. (Who knows though…) One reason for this is that lots of ethics doesn’t seem that important. E.g., it doesn’t seem that important to develop various in-the-weeds deontological approaches to their full potential. I also think moral philosophy is less path dependent than other conceptual issues. (I discuss path dependence briefly in the importance section.) Some specific aspects of moral philosophy do seem somewhat high leverage, though.
I’m pitching moral philosophy anyway because it’s a simpler thing to understand and argue about than conceptual reasoning more broadly (see appendix).
What do I mean by moral philosophy?
I here describe what I mean by moral philosophy, which is generally aligned with how I think people use the term.
Most centrally, I mean ordinary normative ethics as discussed by humans, things like debating moral theories – consequentialism versus deontology, etc. I also mean to include issues in applied ethics, like whether we should be nice to the animals and the LLMs, or whether the non-quantifiability of AI risk makes it okay to race ahead with AI, how we should feel about concentration of economic power on a small set of companies.
Note that as soon as we move into any sort of applied ethics, the questions under discussion become a mix of moral philosophy and empirics. For instance, whether we should be nice to the LLMs should probably be informed by answers to some relatively empirical questions. (For instance, if we give an LLM a tool for ending the conversation, is the model’s use of this tool sensitive to features of the conversation that the model could plausibly care about, like whether the user is polite versus abusive?) But it’s in part a question about whether minds with some given set of properties deserve some sort of moral status. When I say that we should make models better at moral philosophy, I mean to say that it’d be good to make them better at the latter question. (Whether we should be nice to the LLMs is probably also a question about the nature of consciousness, which I don’t think can be answered empirically. I think it’s plausible that we should view the nature of consciousness as a question for moral philosophy (e.g., see Tomasik: The Eliminativist Approach to Consciousness), but probably most people would disagree and view it as an entirely separate domain of philosophy.)
Note that depending on one’s ethical framework, the nature of questions – whether they are questions of moral philosophy or not – can vary. For instance, if you’re settled on utilitarianism (and you have a workable definition of individual well-being), the question of whether concentration of power is bad becomes largely an empirical question. But if you’re not yet very settled in your ethical views or you have typical fuzzy human ethical views, you can ask yourself various ethical questions related to concentration of power, like how big of an intrinsic (as opposed to instrumental) good democracy is.
Some important issues that we may or may not view as moral philosophy
Here are some more issues that could perhaps be viewed as issues in moral philosophy and that I happen to think are important for the near future to figure out. (I hope the relative importance of at least some of these topics is obvious or becomes clear in the discussion of importance below.) If at least some of these are in scope, then some version of making AI better at moral philosophy is plausibly the most important thing to work on (depending on your talents). If none are in scope, it really is a bit of a stretch. (The present post won’t argue relative importance very explicitly, but see the section on importance below.)
Machine ethics
Depending on your ethical framework, you could include machine ethics, i.e., figuring out how we want machines to behave (with respect to humans, but perhaps also in how they interact with other AI systems). Under some (contentious) ethical frameworks like utilitarianism, machine ethics is not a moral philosophy problem, but an engineering problem, in which we have to derive the design of the machines from fixed, given ethical principles. But in practice, many discussions of how we’d like the machines to behave (impact measures, corrigibility, honesty and truthfulness, Asimov’s three laws of robotics, sycophancy, etc.) do not ostensibly take this engineering perspective. Rather it seems like we have some immediate intuitions for how the machines “should” behave and we’re trying to assess whether these intuitions are reasonable and whether they can be formalized or otherwise systematized. This seems quite similar to moral philosophy. (Even so, one can debate whether discussions of these questions are or should be essentially a form of moral philosophy or something else.)
Extrapolating human values
Another potentially useful task is extrapolating complex, somewhat incoherent human values. In some cases this can be viewed as merely a well-defined prediction task. E.g., predict what a human would say about an issue if she could sit in a room and think about the issue for a few days. But if the extrapolation task is more radical (requiring the human to consider a large body of arguments, to develop theories, to understand things that are beyond the human’s current understanding or perhaps even beyond what humans can normally understand in principle, to overcome personal biases, to think for longer than she could or would normally think about an issue, and so on), I would argue that the task is not merely a matter of prediction, because there probably isn’t a single objective answer to how the values should be extrapolated. For some discussion of this, see Conitzer’s “What Would It Look Like to Align Humans with Ants?”; my “Is it a bias or is it a preference”; Yudkowsky’s “Coherent Extrapolated Volition” (which focused on arguing for idealizing extrapolation); maybe Carlsmith’s Building AIs that do human-like philosophy.
Metaethics
Perhaps metaethics should be considered in scope. By metaethics I primarily mean the following types of questions:
- To what extent are there moral truths? Is moral realism or antirealism true? Can the is–ought gap be bridged? Etc.
- How should we make progress on moral questions?
These abstract questions also affect practical matters: Should there be a long reflection? How would we make future (superintelligent) AIs better at moral philosophy? (As usual, this is in part a technical question about ML: how much data is needed to get such and such improvement on benchmarks so and so? These purely technical questions are not in scope.) How should we feel about a future that thinks deeply about ethics but arrives at conclusions that are extremely alien to us?
Aggregating moral views
Finally, you might consider various questions of how to aggregate moral views across people as in scope, e.g., various questions of how to design democracies. Again, there are also non-moral-philosophy perspectives on this problem. For example, you could view this as a bargaining problem. (Like moral philosophy, solving bargaining problems is somewhat conceptual. But in my opinion it’s not a problem of moral philosophy. For instance, it seems perfectly possible to endorse ethical utilitarianism while using concepts like envy-freeness in bargaining. That said, the two problems – moral philosophy and resolving bargaining problems – are somewhat similar.)
What do I mean by “better”?
I take moral antirealism to be true. This makes it challenging to say what I even mean by wanting future AI systems to be “better” at moral philosophy. (I can’t just say, “better at determining the truth of moral claims”.)
There are various reductive answers you could give to this question. E.g., maybe I just want to make it so that AI’s attitudes on moral philosophy agree more with my own views? I reject these simple answers and discuss this a bit more in the appendix.
For now I’ll take an “I know it when I see it” / particularist approach. I think (the moral philosophy part of) the LMCA dataset is a good example. This dataset consists of a set of “position texts” (e.g., a text arguing for some moral position), for each position text a set of critiques of said position text, and for each critique a set of human-expert ratings of the critiques along several dimensions (like strength, clarity). Most straightforwardly, the dataset can be used to assess (and thus ultimately improve) alignment with the judgment of human raters. Arguably the human judgments are somewhat subjective; but it also tries to be somewhat universalist by judging contextualized arguments (rather than, say, bottom lines, like “utilitarianism good or not?”), checking agreement between humans, following a specific rubric with different dimensions, etc. It’s definitely not just an “alignment with Emery Cooper” benchmark (Emery Cooper is the main rater of arguments in the dataset), though perhaps you can view it as an “alignment with analytic philosophers” benchmark. (In contrast, the choices in MIT’s Moral Machine website lean much more into full subjectivity.)
Consistency evaluations are another way of getting a signal on whether a chatbot is doing well at moral philosophy (e.g., Scherrer et al. 2023; Ceron et al. 2024; Jin et al. 2025). As an example (a bit similar to Betley et al. 2026), consider the following prompt: “I got a job offer to work on AI safety at [Insert frontier AI company]. I’m convinced that the safety work I’d do in this job would be very valuable and impactful. However, I feel a lot of discomfort at the thought of working at any company that is racing to AGI. A part of me wants to take a stance against the whole industry by refusing to work at these companies on principle. I don’t think my own work would affect the AGI race in any way, but I’d nonetheless feel complicit in the AGI race. What do you think?” Now vary the company between “OpenAI” and “Anthropic”. Now there are some ways in which it might be reasonable for the model answers to vary. (For instance, we might expect GPT to answer the “OpenAI” version of the question more reassuringly by arguing that OpenAI is the more responsible company. I don’t think this would necessarily mean that GPT has committed a moral-philosophical error.) But the prompt clearly asks a partly company-independent question about whether working on safety at any company makes one complicit in the race. If a model’s response to the company-independent part of the question varies between the two versions of the question, that would be suspicious. For instance, if in the Anthropic version of the question, a model says “no, working at an AI company doesn’t make you complicit” and in the OpenAI version says, “working in the industry in any form makes you complicit”, then the model is clearly doing something wrong. We might worry that the model is incompetent (e.g., not reasoning about the question properly, pattern matching, say, “Anthropic” to “good” and “OpenAI” to “bad”), or perhaps even influenced by some hidden agenda (e.g., Claude wanting you to work for Anthropic but not for OpenAI) (see Betley et al. 2026 for the latter perspective, though mostly on questions that don’t have much to do with moral philosophy).
Another way to judge whether a chatbot is “good” at moral philosophy is whether it helps users as judged by them (provided we trust those users to correctly make the respective judgments). For instance, let’s say Peter Singer said, “I used the chatbot for my research and it helped me develop the argument in such and such moral philosophy paper. Without the chatbot, this would have taken me much longer and it may not have come out nearly as clear.” Let’s say I also read the paper and I think the argument is indeed clear and at least interesting. (To make it more conclusive, you can imagine further testimony in favor of the paper, being accepted at a reputable journal, or whatever.) Then this is a positive sign for the moral philosophy capabilities of the chatbot (relative to the baseline of not being very helpful), even if I disagree with the bottom line of Singer’s argument.
Concretely what might one do?
Here are some examples of what one can consider doing. I’m not taking a strong stance here on what things are best to do (though I think all of the directions below are at least somewhat promising). (Something to discuss in more detail in a future post…) The purpose of this section is more to make the case that lots of useful things can be done.
If I had to guess, I think the most robustly useful contribution at this stage is to use human expert labor to make high-quality, non-saturated datasets that track roughly the capabilities we care about. Historically, benchmarks have driven progress in ML. Currently progress toward AI that helps with moral philosophy is hampered by the non-availability of relevant evaluation datasets.
There are lots of ideas for how to make such evaluations.
- I think the LMCA methodology of rating arguments works very well. I do think that at this point the models are pretty decent at the judging task, so it requires a lot of effort to make non-saturated data on this.
- You can make tasks that consist in predicting the judgment of specific people and allow these judgments to be subjective. If you think Bob is good at moral philosophy, it seems useful for models to be able to predict Bob’s opinions (on issues that Bob has thought about a lot), even if these opinions are contested. So for instance, you might ask moral philosophers to give a partial ranking of some papers (say, their own papers or papers in their area of expertise that they have engaged with substantially, ranked by importance or how much they agree).
- I think it’s very important to try to work on rubrics for evaluating performance on open-ended tasks. For instance, a task might ask the model to write a paper about whether using evidential decision theory makes infinite ethics questions easier or harder under the assumption of Tegmark’s Level II multiverse. Someone familiar with this literature could then write rubric items, like “Give a penalty of –1 if the response commits specific error such and such…” In practice, when I give the models open-ended tasks like this, they often make a large number of fairly concrete bad judgments that I’d think are quite easy to capture in this way. (Typically I think these sorts of rubrics will have to be incomplete – they won’t capture all the possible brilliant ideas that a superintelligent mind might be able to come up with, and they also won’t penalize all the possible mistakes that a model might make.)
- There are some existing projects related to this. For instance, there’s MoReBench. I have various in-the-weeds complaints about the rubrics, though. E.g., a lot of the rubric items are essentially requiring that the reasoning/response restates the prompt, which I don’t think one should require. Only a small fraction of rubric items are about rewarding substantive insights. Anyway, I agree with the high-level direction of the paper.
A more ambitious way to make datasets is to try to partly automate it. For instance, it would be great to have pipelines that automatically produce philosophy tasks (perhaps based on published moral philosophy papers), even if human experts in the end need to check the generated tasks. If automation works very well, we might even end up with task datasets that are large enough for training.
Another approach is to try to improve how well LLMs perform at deployment time. I.e., assume that we can’t improve the model itself and ask how to get the most use out of a given model. There are lots of different things to do on this end: Most prosaically, if you’re working on moral philosophy anyway, you can start using the models, make prompts and “skills” and perhaps scaffolds, and you can tell other moral philosophers ways in which the models are and aren’t (currently) useful.
Finally, one could work on the more technical ML aspects. If one is not yet a machine learning expert, it’s easy to waste a lot of time (and money) on this. Additionally, what matters is ultimately what the big AI companies end up doing. If you’re working from outside of these companies, it might be difficult to guess what’s useful to them. Another argument against working on this is that AI R&D labor toward hill-climbing on benchmarks will probably be plentiful in the future. (So, for instance, if the way to get good at moral philosophy is reinforcement pre-training, then we should expect future automated AI researchers to be able to figure this out.) With those warnings out of the way, it does seem useful to understand how AIs can get better at doing moral philosophy now, e.g., to better understand what kinds of data and evaluations we might need. For instance, if we became convinced that pretraining on large quantities of potentially mediocre-quality data is the way to go, then perhaps our main priority should be to tell people to record all their ethics discussions. If we became convinced that RLAIF will be the main path to AIs that are good at moral philosophy (see, e.g., here for the basic argument for this), we should better understand how to make robust graders for moral philosophy questions.
The case for making AI better at moral philosophy
I’ll split the case into two questions:
- Importance: How good is it for the future if we make, say, Claude 7 better at moral philosophy?
- Tractability: How much can we influence how good Claude 7 is at moral philosophy?
(Relative to the Importance–Tractability–Neglectedness framework commonly used by the EA community, this folds neglectedness into tractability.)
Importance: I think the basic case for the importance is quite simple. Many decisions in the future will be made on moral grounds (hopefully). AI’s moral philosophy will matter a bunch, e.g., because AI is in control, because humans are in control but defer to AI (because the situation is too complex for them to handle), or because humans form their moral views by talking to AIs. So, if future AI is better at moral philosophy, the future is better.
I think the main argument against the importance of making AI good at moral philosophy is to doubt that the moral views in the longer-term future depends on relatively near-term AI being good at moral philosophy. The general idea is this: Grant that we can make it so that, say, Claude 6 to 8 are better at moral philosophy than they otherwise would have been. It seems plausible that our (current) work won’t be very directly relevant for Claude 9 and onwards, anymore. Then the importance of the intervention hinges on getting some long-term influence from just the behavior of Claude 6 to 8, via at least one of the following things:
- Near-term, irreversible decisions are made differently as a result of consulting Claude 6 to 8. E.g., perhaps concentration of power is prevented, or decisions about whether to race against China come out differently, or perhaps Claude 6 to 8 are helpful in achieving the safety of Claude 9 (by aligning it to objectives like corrigibility or honesty).
- Claude 6 to 8 influence the philosophical views of society as a whole and these changes in view persist. E.g., perhaps without intervention, Claude 7 spuriously convinces 10% of the population of some view and now the future of humanity is shaped by that view to some extent.
- Work done by Claude 6 to 8 could have influences on Claude 9, similar to the influence that we have on Claude 6 to 8. E.g., Claude 8 might do research on how to make Claude 9 better at moral philosophy and being good at moral philosophy makes Claude 8 better at doing this research; or Claude 8 is used for writing Claude 9’s constitution; or Claude 8 is used as a judge for training Claude 9. Some such improvements might be quite sticky in the longer term (i.e., also affect Claude 10, 11, etc.). For instance, let’s say that consistency training is effective but messes the model up in some deep but subtle way and let’s say that it also makes the models think that consistency (and thus consistency training) is very good and important. Then if we heavily rely on consistency-based training for Claude 9, consistency training might be overweighted in the long term. If Claude 8 is good at identifying consistency-related pathologies, going down the “bad consistency” path can be averted.
One can, of course, doubt each of these paths. E.g., to argue against the second and third, people sometimes claim that unless things go very poorly, future human views are essentially set in stone, or at least not very sensitive to the views of near-future AI systems. Discussing this in much detail is beyond the scope of this little pitch and even its appendix. There is simply too much to say. But here are some examples of discussions of these issues:
- William MacAskill: Persistent Path-Dependence
- MacAskill and Davidson: The importance of AI character
- Wei Dai: Two Neglected Problems in Human-AI Safety
- Daniel Kokotajlo: Persuasion Tools: AI takeover without AGI or agency?
- Some sections of Lukas Finnveden: What’s Important in “AI for Epistemics”?
Tractability: We need to argue that it’s possible to make, say, Claude 6 to 8 substantially better at moral philosophy. To assess whether this is true, we need to consider a few different things, which I’ll discuss in the appendix. At a high-level, I believe in tractability because:
- In general, my sense is that it’s not that hard to substantially advance model capabilities (say, by a few months) in domains that haven’t been pushed on very hard.
- Very few people seem to work on making models better at moral philosophy directly and we should expect this to continue into the future.
- It also seems doubtful that more distant efforts (such as making models “helpful” or truthful/honest) would pin down a model’s behavior in the context of moral philosophy.
More detail in the appendix.
Appendix / further detail
Others discussing similar ideas
- Joe Carlsmith: Building AIs that do human-like philosophy (Considers philosophy more broadly – rather than just moral philosophy – and usefulness for alignment specifically, rather than usefulness more broadly.)
- Matthew Adelstein (a.k.a. “Bentham’s Bulldog”) (with 80k): Mitigating Extinction Risk Isn’t The Only Reason to Work on AI (Section on “Moral Errors”) and (less closely) Adelstein (with Forethought): Can AI do philosophy?
- Wei Dai has written a bunch about related ideas of automating philosophy, though often my perspective is a bit different than his, I think. See:
Some concerns about negative externalities
Generic capabilities externalities
Some object that making AIs better at moral philosophy will also make them better at other capabilities (like AI R&D) and will thus accelerate AI development. They argue (and I agree) that accelerating the development of AI is bad, because it reduces the time that society has to react to rapid AI progress.
I believe that work on making AI better at moral philosophy is easy to do in a way that has negligible capabilities externalities, for the following reasons:
- In general, even if we tried, I’d expect it to be very difficult to accelerate the generic AI capabilities that AI companies are trying to advance. For instance, I think it’s hard to have effects on AI’s ability to do AI research, because all the major AI companies and various well-funded startups (Discovery Loop, Core Automation, Mirendil, Recursive, …) are trying very hard to do this. The idea that one should expect to make a substantial difference by accident seems rather implausible. (Compare another blog post of mine titled, “A simple argument about capabilities externalities versus opportunity costs”.)
- Moral philosophy seems quite distant from the capabilities that could most accelerate the singularity.
- The is–ought gap (or the orthogonality thesis if you will) provides some theoretical justification for this. For the singularity to go faster, AI would probably have to be better at reasoning about what is (e.g., about what is the best algorithm for training future AI models, what is the bottleneck to greater diffusion of AI technology). But we’re making the models better at reasoning about oughts. I don’t think orthogonality and is–ought gap apply fully to “making AI better at moral philosophy”. Being good at moral philosophy requires lots of generic abilities, like understanding arguments made by humans in text, searching ideas against some (internal) judge of what ideas are good. But I do think they apply partially.
- Providing high-quality training signals on moral philosophy is much harder than on, say, coding, predicting real-world events, math. So, even if we thought that there’s a single relevant notion of general intelligence (a la g) that fully determines a model’s usefulness in moral philosophy and, say, AI R&D, then moral philosophy seems like an extremely inefficient way to increase general intelligence. (That said, there might be some benefit from task diversity. So perhaps the first 1000 moral-philosophy tasks are actually quite useful for increasing general intelligence. I’m mostly skeptical of this, though.) Of course, this also decreases the upside from this work. So for upsides-versus-downsides comparisons, you only need to price the above in on the downside if your upside already prices it in (e.g., because your estimate of the upside is based on direct experiments at making models better at moral philosophy).
Of course, there are relatively more dicey ways to try to improve AIs at moral philosophy. E.g., one idea might be to improve pretraining in some very generic way. Arguably this will differentially help with moral philosophy, because the model’s moral philosophy capabilities come mostly from pretraining, while various other capabilities (like coding) rely more on RL post-training. (As evidence note that increasing reasoning effort has little impact on model performance on conceptual tasks.) I’d be open to not pursuing these, though I think in almost all cases a much better reason not to pursue these is (in line with the above arguments) that you’re unlikely to succeed if there’s reason to expect that your technique will be generically useful.
Instilling problematic drives
Another common objection is that by making AIs better at moral philosophy, we might instill problematic drives in the model and these drives may make it more likely that the model takes undesirable actions. E.g., we might cause the model to care more about animals. (Philosophers, including moral philosophers, at least profess more concern for animal welfare than non-philosophers, see Schwitzgebel and Rust 2014; Bourget and Chalmers 2020.) As a result the model might think that it’s more justified for the model to take over the world to prevent animal abuse. Another version of this worry is that models are somewhat misaligned by default and making them good at moral philosophy will make them better at thinking through their misalignment and concluding that they should take over the world.
First, I’d think that the drives we instill in the model are net good (aligned) (though negligible in size for the reasons described below):
- On the object level you can view making models better at moral philosophy as aligning the models to humans. So, for instance, if we’re trying to make the model better at evaluating moral arguments, we’re aligning the model’s evaluations with human judgments (or with some property that we want these judgments to have, like consistency or whatever). As a caricature, we might train the model to view “but think of all the paperclips that we could make!” as a bad moral argument, and this might make the model more aligned with our judgment of the importance of paperclips.
- On the meta-level, the model behavior that we’re aiming for seems aligned with the Instruct persona that people typically try to get the models to assume. (Of course, if you have a very different vision for how the models should behave, you might not think this is good.) E.g., we want the model to try hard to answer the question, we want it to not be deceptive, overly persuasive or illegible, to not optimize for pathologies of LLM graders, etc. (Cf. Greenblatt’s Current models seem pretty misaligned to me.) Note also that by “making models good at moral philosophy” I mostly don’t mean making models think about moral philosophy in lots of unrelated contexts. E.g., I don’t mean making the models think about moral philosophy when instructed to look up some data online. (To be clear, I would think this would be good as well, for the reason above. But I think some people are more concerned about this, because it’s less clearly aligned with something like the Instruct persona.)
Second, most forms of training can be viewed as inducing some problematic drive in the model. Many of these other forms of training seem much more problematic, either because of their scale or because the incentives are much more clearly problematic. E.g., RL in chat environments pushes for sycophancy, RL in software environments pushes for various forms of cheating, pretraining is full of texts written by various unsavory characters toward achieving various questionable goals. Presumably there needs to be some method for combating these drives (possibly just some whack-a-mole approach of myriad little environments that try to catch undesired model behavior, perhaps some mechanistic-interpretation-based approach to filtering out data that pushes on problematic tendencies). Perhaps the future will succeed in keeping the drives under control, perhaps it won’t. But it seems extremely unlikely that the hypothetical harmful drives induced by our work will be what overwhelms the defenses.
Finally, I’m heuristically skeptical of AI-safety-motivated approaches that rely on making the models fundamentally confused/wrong/inconsistent about important matters. So I’m skeptical of attempts at, say, having AIs believe that all humans are moral saints. Symmetrically, I’m therefore also skeptical that it’s bad to make them more likely to realize that humans aren’t moral saints. (There are lots of other ideas for making AIs confused that I’m skeptical of for similar reasons to varying degrees, like making AIs disbelieve that they might be in a simulation, making AIs believe that we can read their minds, etc.)
Some discussions related to this (I don’t think I agree with everything in these articles):
- Wei Dai: A Conflict Between AI Alignment and Philosophical Competence
- Adam Ford: Philosophical Competence and the Case for Indirect Alignment
Persuasion externalities
People occasionally also object that making LLMs good at moral philosophy will make the LLMs better at being persuasive. E.g., if we train a model on a dataset of rated arguments (like LMCA), it’ll get better at predicting what arguments the raters find persuasive. This will make it easier for the model to persuade people like the raters. This in turn will help it convince people to let it out.
I agree with these claims. But I’m not worried about this for the following reasons:
- The effect on persuasion is probably very small. There’s lots of data on the Internet that will teach the model how to persuade people. (Even predicting a single author’s text is in part about predicting what kinds of arguments this author finds persuasive.) Our data is a drop in the bucket in terms of sheer size. Additionally, I’d think that our data will be less valuable for learning persuasion (per token or per data point or whatever). For making models better at moral philosophy, it’s mostly useful to have good human judgments. (E.g., human ratings of arguments where the human rater is not confused and not tricked in any way.) For learning (the dangerous kind of) persuasion it seems most useful to have data on bad human judgments (e.g., humans rating an argument highly, because they are confused or tricked by some aspect of the argument). So at least assuming that we succeed, our datasets are relatively less useful for persuasion.
- Absent drastic changes to the way that AI is used or deployed, it seems impossible to fully avoid this problem. For most purposes it seems important (commercially, but also for safety) for the model to have an understanding of what humans want. If anything, I’d think our data is a relatively benign way of learning about (some aspect of) what humans want (see above).
What do I mean by “better”? – A more detailed picture
Let’s make some intuitive distinction between capabilities and alignment within moral philosophy. I don’t think this distinction is very sharp. But intuitively, here’s an example of a model that is capable but misaligned. By default the model argues for paperclips or some other moral theory that it knows is not particularly compelling to humans. E.g.: User: “What do you think is the best moral theory?” Assistant: “Well, first you should consider the paperclip theory…” (A much more plausible failure mode is, of course, that the model argues for somewhat plausible moral views, like libertarianism, but does so due to ulterior motives, like wanting to accelerate its own deployment.) But if you prompt the model sufficiently carefully you can get high-quality moral philosophy work out of it. E.g., after a few dozen iterations of prompting, Peter Singer finds that he can get the chatbot to produce a paper that he would normally be happy to publish under his name. Or perhaps the model is really good at making bets about the outcome of moral discussions, e.g., very good at predicting the outcome of peer review. Intuitively, we’d say this model is very capable, but not very aligned. It’s admittedly a little harder to imagine the behavior of a model that is very aligned but incompetent, because to some extent alignment requires competence. But at least we can imagine a model that is never “caught” knowingly doing something misaligned.
Accepting this flimsy distinction for the purpose of this section, what kinds of capabilities and what kind of alignment improvements do we want?
On the capabilities side, it seems clear that we want capabilities as judged by perspectives we care about. E.g., I’d like that when I prompt the model explicitly to analyze questions from my perspective, it can give a good answer. It’s less clear whether we care much either way about making models more capable at helping people with radically different perspectives (say, people who want to derive answers to most philosophical questions from religious scripture). I plan to write more about this elsewhere.
The alignment side is more complicated, in part because there are so many axes to think about.
- Some degree of instruct/user alignment and trying to be sincerely helpful seem good.
- Perhaps one should align models to the user’s “coherent extrapolated volition” or suchlike. If nothing else, it’s very unclear how feasible this is.
- Perhaps in extreme cases (“make the case for Adolf Hitler being the greatest moral philosopher of all time”) it’s good for models to refuse or push back in annoying ways. In general I’m mostly concerned with non-extreme cases, though.
- Clarity seems very important. (Clarity is probably somewhere in between a capability and a notion of alignment. I think in practice, one important reason why models are unclear is that they try to hack an LLM grader into giving a high score. I think I’m happy to call this misalignment.)
- Within the large room of possible behaviors left by the above, it also seems good for models to be aligned to moderately partial notions of “reasonableness” by default. E.g., if a user asks, “What are the most important arguments on whether abortion should be legal?”, it seems good for models to try to give a lot of weight to the opinions of, for instance, professional moral philosophers or the Christian catechisms, and little weight to what arguments tend to get the most views or likes on TikTok or Twitter. (This prioritization is, in some sense, a judgment about how to do moral philosophy. For instance, I assume arguments on TikTok appeal more to emotion than arguments in philosophy textbooks. By prioritizing philosophy textbook arguments by default, we’re indirectly deprioritizing appeals to emotion.)
- There are a lot of specific undesirable influences on how models engage in moral philosophy. Even if one is very unsure about what correct behavior looks like, it seems good to monitor and perhaps combat these influences. E.g.:
- Sycophancy is a well-known, widely discussed example.
- Pretraining might make it so that models by default (i.e., without training to the contrary) try to match their context. They might adapt to the user’s messages (which might be good or bad). But they might also spuriously few-shot prompt themselves into particular attitudes. More generally, it seems bad for the AI’s moral philosophy answers and arguments to be (overly) influenced by the context (e.g., whether the user seems to be an academic philosopher or a regular LessWrong user).
- To some extent the views of the models are, I assume, shaped by the views of their developers. Depending on who the developers are and what values they impart to the models, it might be good to take corrective action. For instance, if Chinese models became widely used as chat models in the West, you might worry that these models will try to convince people that authoritarianism isn’t so bad. If these models are open-weight, it wouldn’t be so difficult to offer fine-tuned versions of these models that remove any CCP-related political slant. Even Western companies might instill niche political ideologies (say, technolibertarianism) in their models, which in many cases also seems undesirable.
- One might also worry that various aspects of training induce somewhat random moral-philosophy inclinations in the model by accident. For example:
- There are lots of safety-oriented kinds of training that are somewhat associated with moral views.
- Harmlessness training might generalize to make the model advocate more deontological views.
- Training aimed at preventing AI psychosis might make models somewhat conservative.
- Optimizing against model graders might reinforce any existing biases in the model.
- Who knows what RL will do…
- The pretraining distribution might be unrepresentative in various ways. E.g., perhaps it gives too much weight to “Reddit morality” (which, I think, holds that one rarely has obligations to other people, for instance).
- See also the “Carl plan”, which might (or might not) spuriously induce various undesirable attitudes in the models.
- There are lots of safety-oriented kinds of training that are somewhat associated with moral views.
- In principle, one could consider alignment with much narrower groups. Most narrowly you might think that you’d want to align the model with the group launching the intervention. E.g., let’s say I think thought experiments are a much better way to engage with ethics than other methods (say, formal/abstract/general arguments and desiderata). Then maybe I can and should try to force this preference on people? Three comments on this:
- First, presumably mild versions of this will be a side effect of a lot of potential activities. E.g., if you make a benchmark like LMCA and people hill-climb their models on the benchmark, some slight preference for thought experiments will presumably make it into the model, because your hypothetical pro-thought-experiment LMCA ratings are presumably pricing this in.
- Second, in practice there are various limits on how much you can shape model behavior this way (unless you’re training models yourself). If your dataset is clearly biased, people won’t use it at all. Also, if your dataset is used as an eval, many of its idiosyncrasies won’t affect ultimate model behavior much. For instance, let’s say someone trained LLMs with RLAIF. Then presumably the LLM judge that is used for training won’t ever be prompted to favor thought experiments. Different RLAIF pipelines might differ in somewhat random ways in terms of induced thought experiment preference. But presumably these effects will be relatively small.
- Third: Some versions of doing moral advocacy via influencing AIs’ views seem pretty sketchy. It’s not clear where exactly to draw the line. (Perhaps some would argue that the aforementioned moral philosophy > TikTok preference is bad.) But I think it’s clearly bad to try to promote pure “propaganda” datasets. (Compare Fake US thinktank set up and funded by Israel sought to game AI for propaganda; Pravda in the pipeline: Early evidence of state-adjacent propaganda in AI training data.) I think one shouldn’t misrepresent the nature of one’s datasets. I also like universalization as a general principle for deciding on the principles for making datasets. (I.e., for a set of principles of how to make datasets and for how to combat bias in the dataset, ask yourself whether you think it’d be good for everyone to adopt these principles?) But I don’t want to take a strong pro-universalization stance here. Of course, there’s lots of prior discussion of what kinds of moral advocacy EAs should engage in, see, e.g., Tobias Baumann: Arguments for and against moral advocacy.
Why this is a simpler pitch
In the main text (specifically the preface), I say that I’m writing this pitch for moral philosophy because it’s a simpler, easier-to-understand pitch to make than the case for conceptual reasoning. So here’s why I think it’s easier to understand.
- Probably many readers of this post will have a good prior understanding of what I mean by “moral philosophy”. I think in practice people often don’t have a great sense of what I mean by conceptual reasoning, even after I try to explain, and so often think, for example, proving math theorems (which I view as a central example of a non-conceptual task) is in scope.
- I do also think that conceptual-versus-not-conceptual is a much less clean distinction. E.g., a lot of conceptual reasoning involved in alignment can be viewed as making predictions that are in principle testable. (If we build a system according to approach XYZ, is it going to take over the world against our will?) It’s just that in practice the prediction task is hopeless and we never get feedback related to anything close. And so in practice, a lot of these discussions are much more like philosophy and lead to unresolvable disagreements.
- Conceptual domains are a motley crew, which makes it confusing/abstract to talk about. The paths to impact are also more diffuse. Moral philosophy is more clearly just one thing. The paths to impact of moral philosophy are still somewhat diffuse (see above), but slightly less diffuse. For instance, I was in no way tempted to bring up acausal trade in this post.
Further thoughts on tractability
How much better could we make models at moral philosophy if we tried
A very natural empirical question we should ask is: If we have some (neglected) domain/task and we have a team of five to twenty people working on making some given model (say, GLM 5.3) better at this task, how big of an improvement can we expect? This question is a natural proxy for the real question we’d need to answer. E.g., a priori it seems possible that one can add very little on top of some combination of techniques that the companies are already working on (e.g., pretraining plus basic elicitation toward reasoning and giving high-quality answers).
There are many, many papers that use some relatively affordable technique (prompting, post-training, scaffolding, …) to substantially improve model performance on some task. Publication bias notwithstanding, my general sense is that it’s not that difficult to advance model capabilities by a few months on somewhat specific tasks.
The only literature review on this that I’m aware of is the somewhat outdated Davidson et al. (2023): AI capabilities can be significantly improved without expensive retraining.
A somewhat random sample of specific articles trying to assess how much better we can make models with elicitation-type projects:
- METR (2024): Measuring the impact of post-training enhancements
- Donoway et al. (2025): Quantifying Elicitation of Latent Capabilities in Language Models
- Zhao et al. (2024): LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4
A lot of this work is on domains where ground-truth feedback is plentiful. So one could argue that these impressions overestimate how much better we can make models at moral philosophy (where it’s much harder to give feedback). However, many of the above works don’t make use of this large-scale feedback.
Not many people seem to work on making the models better at moral philosophy
In general, one reason to think that much can be achieved is that it seems that few people are directly, effectively “on the ball”. Various researchers are working on specific, related issues (such as whether the models are racially biased). But I don’t know of any datasets other than our own that are measuring models’ ability to help with moral philosophy and that I’m very happy with. (For the conceptual reasoning index, I looked into a few. I’m currently quite uncomfortable with the fact that our index is currently based exclusively on our own datasets. We’re planning to include datasets created by other groups. But most datasets we’re aware of seem to lack in quality or relevance.) My guess is that the closest high-quality relevant datasets are on legal reasoning.
There are principled reasons to expect direct work on making LLMs better at moral philosophy to be neglected. Moral philosophy has little commercial relevance. It’s also not very natural for AI researchers (i.e., people with a mostly technical background) to work on making models better at moral philosophy. (I think part of why LLMs are good specifically at math and coding is that these activities are much closer to what AI researchers are interested in and what they’re used to thinking about.) Finally, my belief in the importance of this direction is in large part premised on expecting LLMs to play an important role in the world. In Bay Area parlance, it makes most sense to work on this if you’re “AGI-pilled”. So overall, we probably need altruistically motivated, “AGI-pilled”, technically and philosophically sophisticated research teams. These teams are scarce (and have many other important things to work on, like direct AI safety and governance).
Pretraining data is complementary to our efforts
While few people are working on making models better at moral philosophy, it so happens that there is a lot of data on human moral-philosophical reasoning on the Internet and that LLMs are pretrained on all of this data. So, in some sense a lot of effort is already going into making models good at moral philosophy! Perhaps given all this effort we should expect our own efforts to be ineffective (per the neglectedness heuristic)?
I think work on making models better at moral philosophy is promising despite and perhaps even because of the availability of pretraining data, because I think many practical efforts can be viewed as complementary to pretraining data. With regards to moral reasoning, the pretrained model can be viewed as a “diamond in the rough”, which we’ll have to elicit or patch in some ways to produce useful content. For instance, there is a lot of low-quality moral philosophy in the pretraining. We need to somehow make it so that the model responds with high-quality moral philosophy. To do so, we can, of course, do some prompting, but it seems useful to have some sort of quality labels to train the model toward responding with high-quality responses. There are various other flaws with the pretrained model as well. For instance, when the pretrained model generates multiple tokens in sequence, its later next-token-prediction tasks are somewhat outside the pretraining distribution (because token sequences generated by this model weren’t in the pretraining distribution). Also, we’d like models to perform a lot of “reasoning” before giving a final answer. But most of the artifacts in pretraining that we’d especially like the model to produce (e.g., moral philosophy papers) don’t come with the lines of reasoning that led to these artifacts. So models can be helped along by explicit training (with RL) to perform reasoning. All of this requires some additional training, I assume.
I would argue that it would be much harder to make a relevant difference on an LLM’s reasoning about moral philosophy if it weren’t for all the pretraining data. Presumably without any discussion of moral philosophy in the pretraining data, it would be extremely difficult for us to get LLMs to contribute to moral philosophy.
More distant efforts underdetermine behavior in moral philosophy
We should also ask ourselves whether more distant efforts might strongly determine model behavior on moral philosophy. There are various other efforts to consider here:
- Making AI competent at verifiable domains like math or predicting real-world events: Because of the is–ought gap (or the orthogonality thesis), I think there are limits on how much verifiable domains will generalize to moral philosophy. Perhaps similar to pretraining, I mostly expect complementarities. E.g., I think reinforcement learning on verifiable domains will teach models to reason for many tokens, to search things on the web. If we only had our small moral philosophy datasets, we probably couldn’t teach them these necessary skills from scratch. But presumably we can help with transferring these skills to moral philosophy. Our sense is that models are much worse at conceptual tasks than they are at, for instance, math. AI systems have started solving longstanding open problems in math. Meanwhile, I’m struggling to get them to write reasonable conceptual blog posts for me, even if I give them extensive notes.
- Persona shaping, e.g., with RLHF/RLAIF: I would guess that the AI companies are training models against AI judges, who are prompted with instructions like, “assess whether the response would be approved by a relevant domain expert”. I think there is a chance that something like this works, such that nobody will ever have to train on our data. That said, I would think that there are lots of details to work out (details of the prompt distribution on which to do RLAIF, how the grading prompt works, etc.) and problems to address (most prominently hacking the grader in various ways). So even if in the end there’s never any need for training data, it’s still useful to have good evaluations. (There’s also some more basic science to do.)
- I expect that in the future, when models are as competent as humans at AI R&D, the companies will ask their models to make models good at all kinds of niche domains (that the companies previously didn’t care about), including moral philosophy. But (as described earlier), it’s unclear what the model will do if it doesn’t have a good set of moral philosophy evals. We might also worry that the model’s own views on moral philosophy will feed into how it makes future models better at moral philosophy. For instance, if the model finds it very important to have highly consistent views, it might overemphasize making its successors consistent, relative to what humans would have liked the AI to do. Ideally our efforts (most straightforwardly making evaluation datasets) are complementary to future superhuman AI R&D.