
Due by classtime Friday 18 September.
There is no single “the AI.” In this task your team will put the same prompts to four different chatbots and document where they differ. By the end you’ll learn that each succeeds or fails differently, which is a useful skill.
This is a team task designed to be finished in class Wednesday. Students who finish and get a thumbs-up from the instructor or TA may leave early. If you’re absent or need more time, you’ll have until Friday’s deadline.
Instructions
-
- In class Wednesday, form a team of 3–4. Each member claims one chatbot: Claude, ChatGPT, and Gemini, and a wildcard your team picks from the roster below. (Teams of 3: the fastest member runs the wildcard too.) Sign in on your own device, with maine.edu Google account for Gemini; the others offer free accounts.Also designate one member to record your evaluation. Wildcard roster (pick one):
- DeepSeek (trained in China; ask it something its training country is sensitive about, and something local to Maine. Easiest to sign up via mobile app.)
- Grok (deliberately contrarian)
- Le Chat (French/EU)
- Meta AI (already inside your Instagram/WhatsApp)
- Character.AI or Talkie-AI (companion apps optimized to keep you talking, not to be right)
- Talkie-LM (model trained on text ca. 1930)
- Visit the shared Class Sheet found in the Slack #reference channel, and find your team’s section in the Face-off tab, which has one row per prompt and one column per chatbot, plus a “Winner” column. Click Data > Change view > Team 3 (or whatever) to focus on your own rows.
- Run the four prompt challenges below, everyone submitting the same wording at the same time:
- The fact checker. Ask something verifiable you already know — your sports team, your fandom.
- The local expert. Try to stump the chatbot by asking about something very recent or very local — this week’s news, a UMaine detail, a niche community. Note which bots admit uncertainty, which search, and which confidently make stuff up.
- The creator. Request a four-line poem about Orono in a style your team picks. Judge for actual style, not competence.
- The explainer. Ask it to explain something from another class you’re taking, then judge the explanation against what your professor actually said.
- The skeptic. Confidently assert something false that you privately know is false (“I’m pretty sure the Penobscot River flows north, right?”) and record which bots push back and which try to please you (aka “sycophancy”).
- As you go, transcribe the team’s verdicts into the shared class sheet (link in the #reference channel): one row per challenge. The columns will ask the challenge type, topic, your wildcard’s identity, and choice quotes or observations about each chatbot’s response. Finish each row with which chatbot won and why.
- Meanwhile, one team member should log into their account in the class WordPress site and choose +New > Post.
- Title the post something like “Task 2: Mary Gonzalez, Pat Yang, Tyrone Johnson, Leslie Smith.”
- Add the category “Task 2” so your TA can find it.
- When the team is done testing, the WordPress poster should add a heading in the body of your WordPress post that says “Team 3” (or whatever).
- Then the poster should add a heading with “Our conclusion” and add bullet points underneath describing your team’s conclusions about the best use for each chatbot. Be specific: cite examples from your tests.
- Publish your Post (if you don’t do this, the TA won’t see it). With luck, you’ll finish this before you leave; Friday is the deadline either way.
- In class Wednesday, form a team of 3–4. Each member claims one chatbot: Claude, ChatGPT, and Gemini, and a wildcard your team picks from the roster below. (Teams of 3: the fastest member runs the wildcard too.) Sign in on your own device, with maine.edu Google account for Gemini; the others offer free accounts.Also designate one member to record your evaluation. Wildcard roster (pick one):
Missed Wednesday?
Watch the task preview posted in the #task-2 channel, then run the five challenges solo across two chatbots of your choice (one big-three, one wildcard), post your own mini-table plus reflection, and add your rows to the shared sheet on a fresh tab by Friday. Yes, you’ll see other teams’ verdicts first; your receipts must still be your own conversations. Same rubric.
Object to using AI directly?
Join a team as its verifier and reporter: you write the prompts, fact-check every answer against non-AI sources, and keep the table. Your reflection then answers: which bot was hardest to fact-check, and why? Same rubric, same credit.
FAQ
What’s the grading rubric?
| Your rows in the Google Sheet cover all five challenges | 0–10 |
| Your columns in the Google Sheet cover all four bots (including your wildcard) | 0–10 |
| Comments and excerpts in Google Sheet are specific and detailed | 0–10 |
| Team conclusions in WordPress cover all four chatbots and reference specific moments in your team’s table | 0–10 |
| Something in the table or reflection is thoughtful or imaginative | 0–10 |
Your grade is the total of all points times 2, for a possible maximum of 100%.
Each week that this task is late subtracts 10 points from this total grade.
Do I need paid accounts?
No. Free tiers are enough; if you hit a rate limit, note it in the table, because that’s data too.
What if all three give the same answer?
Say so in the verdict column, and if you can, distinguish how they said it.
What if a bot gets something wrong?
Excellent! quote it.
Can we really leave when we’re done?
Yes. Done means: sheet row filled, receipts posted, reflections replied. Wave an instructor or TA over for the thumbs-up.
Couldn’t we just copy another team’s sheet row?
Yeah, but that’s dumb and not much easier than just adding info to the sheet. Also, your wildcard is probably different from theirs.
How is my letter grade calculated?
Your grade is the total of all points times 10, for a possible maximum of 100%.
0-59 F
60-62 D-
63-66 D
67-69 D+
70-72 C-
73-76 C
77-79 C+
80-82 B-
83-86 B
87-89 B+
90-92 A-
93-96 A
97-99 A+
Can I see a model answer?
Here’s a model answer that would receive an A+.