Birdwatch Archive

Birdwatch Note Rating

2024-10-06 13:39:06 UTC - NOT_HELPFUL

Rated by Participant: 9F02E1BA437E9D831E34B7F1F5D256DD09EAA574248D90122D5713CB577E0180
Participant Details

Original Note:

Claude 3.5 Sonnet does not outperform o1 in terms of reasoning using this prompt. https://x.com/ikristoph/status/1842807922214175105?t=7D5ymBDVIErL9hcMRhNYvQ&s=19 Although there are some results, the title is factually a clickbait. Outperforming o1 through prompting alone is virtually impossible. https://openai.com/index/learning-to-reason-with-llms/

All Note Details

Original Tweet

All Information

  • noteId - 1842837832898850925
  • participantId -
  • raterParticipantId - 9F02E1BA437E9D831E34B7F1F5D256DD09EAA574248D90122D5713CB577E0180
  • createdAtMillis - 1728221946382
  • version - 2
  • agree - 0
  • disagree - 0
  • helpful - 0
  • notHelpful - 0
  • helpfulnessLevel - NOT_HELPFUL
  • helpfulOther - 0
  • helpfulInformative - 0
  • helpfulClear - 0
  • helpfulEmpathetic - 0
  • helpfulGoodSources - 0
  • helpfulUniqueContext - 0
  • helpfulAddressesClaim - 0
  • helpfulImportantContext - 0
  • helpfulUnbiasedLanguage - 0
  • notHelpfulOther - 0
  • notHelpfulIncorrect - 0
  • notHelpfulSourcesMissingOrUnreliable - 0
  • notHelpfulOpinionSpeculationOrBias - 0
  • notHelpfulMissingKeyPoints - 0
  • notHelpfulOutdated - 0
  • notHelpfulHardToUnderstand - 0
  • notHelpfulArgumentativeOrBiased - 0
  • notHelpfulOffTopic - 0
  • notHelpfulSpamHarassmentOrAbuse - 0
  • notHelpfulIrrelevantSources - 0
  • notHelpfulOpinionSpeculation - 1
  • notHelpfulNoteNotNeeded - 0
  • ratingsId - 18428378328988509259F02E1BA437E9D831E34B7F1F5D256DD09EAA574248D90122D5713CB577E0180