Facebook’s Civil Rights Audit
Chapter Three: Content Moderation & Enforcement
Content moderation — what content Facebook allows and removes from the platform — continues to be an area of
concern for the civil rights community. While Facebook’s Community Standards prohibit hate speech, harassment,
and attempts to incite violence through the platform, civil rights advocates contend that not only do Facebook’s
policies not go far enough in capturing hateful and harmful content, they also assert that Facebook unevenly
enforces or fails to enforce its own policies against prohibited content. Thus harmful content is left on the platform
for too long. These criticisms have come from a broad swath of the civil rights community, and are especially acute
with respect to content targeting African Americans, Jews, and Muslims—communities which have increasingly
been targeted for on- and off-platform hate and violence.
Given this concern, content moderation was a major focus of the 2019 Audit Report, which described developments
in Facebook’s approach to content moderation, specifically with respect to hate speech. The Auditors focused on
Facebook’s prohibition of explicit praise, support, or representation of white nationalism and white separatism.
The Auditors also worked on a new events policy prohibiting calls to action to bring weapons to places of worship
or to other locations with the intent to intimidate or harass. The prior report made recommendations for further
improvements and commitments from Facebook to make specific changes.
This section provides an update on progress in the areas outlined in the prior report and identifies additional
steps Facebook has taken to address content moderation concerns. It also offers the Auditors’ observations and
recommendations about where Facebook needs to focus further attention and make improvements, and where
Facebook has made devastating errors.
A. Update on Prior Commitments
As context, Facebook identifies hate speech on its platform in two ways: (1) user reporting and (2) proactive
detection using technology. Both are important. As of March 2019, in the last audit report, Facebook reported
that 65% of hate speech that it removed was detected proactively, without having to wait for a user to report it.
With advances in technology, including in artificial intelligence, Facebook reports as of March 2020 that 89% of
removals were identified by its technology before users had to report it. Facebook reports that it removes some posts
automatically, but only when the content is either identical or near-identical to text or images previously removed
by its content review team as violating Community Standards, or where content very closely matches common
attacks that violated policies. Facebook states that automated removal has only recently become possible because
its automated systems have been trained on hundreds of thousands of different examples of violating content and
common attacks. Facebook reports that, in all other cases when its systems proactively detect potential hate speech,
the content is still sent to its review teams to make a final determination. Facebook relies on human reviewers to
assess context (e.g., is the user using hate speech for purposes of condemning it) and also to assess usage nuances in
ways that artificial intelligence cannot.
Facebook made a number of commitments in the 2019 Audit Report about steps it would take in the content
moderation space. An update on those commitments and Facebook’s follow-through is provided below.
42