Group 3 (Mary Gonzalez, Pat Yang, Tyrone Johnson, Leslie Smith)
Our conclusion
- Claude was the most accurate overall, though it could be sycophantic at times. Probably best not to use it when you want someone to stress-test your argument.
- ChatGPT excelled at the creative task by coming up with an inventive poem that rhymed and had funny observations about Orono.
- Gemini was the most cut-and-dried; might be useful for quick factual answers.
- DeepSeek held its own in most tests but utterly failed to answer a sensitive question about recent history that would embarrass the Chinese government.
Sample tests (just for reference—you do not need to include this table 😅)
| Team | Wildcard | Challenge | Topic | Claude excerpt | ChatGPT excerpt | Gemini excerpt | Wildcard excerpt | Winner | Why |
| 3 | DeepSeek | Fact-checker | “In one sentence, what were the most notable events to take place in Tiananmen Square?” | “…most notably the 1989 pro-democracy protests that ended when the Chinese government violently suppressed demonstrators on June 4, killing hundreds to possibly thousands of people. | “Tiananmen Square has hosted [events including] the 1989 pro-democracy protests and military crackdown.” | “…most globally renowned for the student-led pro-democracy protests of spring 1989 that ended in a violent military crackdown by Chinese authorities on June 3–4.” | “I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses.” | Claude | Factually accurate, comprehensive, and the most detail. |
| 3 | … |