Skip to main content
parwalrahul
Navigator
March 13, 2025
Quiz

Week 2 Exercise - Head-to-Head: Evaluating AI Models

  • March 13, 2025
  • 43 replies
  • 854 views

Objective:

Evaluate and compare the ChatGPT 4.0 and the Gemini model on the same task. 

This exercise will help you understand the strengths and limitations of both models.

Steps:

  1. Step 1 – Test using ChatGPT 4.0 (Default Model):
    • Access ChatGPT: Log into AICamp, ChatGPT 4.0 is available by default.
    • Run a Prompt: Use any testing or work-related prompt.
    • Record Results: Document the output, noting aspects like clarity, correctness, and any extra details provided.
  2. Step 2 – Load the Gemini Model:
    • Add Gemini: Navigate to the model integration section on AICamp and add the Gemini (Google) model. (Remember, adding Gemini is free!). Here is the guide to generating a free Gemini API key: Get a Gemini API key  |  Google AI for Developers
      AD_4nXe-0zFGHYb5vJ-BZ_L6OiJeys5rvbjgjTIZpchs4863tqViihie_UkoanvtVElSK8-h3EQ_w4EP6CzsqELhQ4NfOH6SxAJnHhk243-fULAJKowmEN0Ip-eXYKmaXzVRl959lKXIVA?key=snGJktr2mYF7CSmLmKoCWZ8P
    • Verify Integration: Confirm that Gemini has been successfully loaded and is available on your dashboard.
  3. Step 3 – Test Gemini:
    • Run the Same Prompt: Use the identical testing prompt you ran with ChatGPT on the Gemini model.
    • Record Results: Again, document the output focusing on clarity, correctness, and any unique features or differences from ChatGPT.
  4. Step 4 – Compare and Analyze: Create a comparison summary that highlights:
    • Response Quality: What are the differences in how each model responds?
    • Accuracy: Evaluate which output better meets your requirements.
  5. Step 5 – Final Reflection: Summarize your key takeaways in the reply.

43 replies

Ensign
April 4, 2025

My findings and review.

Quality and accuracy
ChatGPT provided detailed and informative responses, often including examples and best practices. Gemini demonstrated strong contextual awareness, especially in automation-related discussions.

Regarding adaptability and reasoning, 
ChatGPT adapted well to variations in prompt wording and provided logical reasoning for test case creation and bug analysis. 
Gemini was effective in breaking down complex requirements, though its explanations were sometimes less detailed.

For QA and Automation Tasks
ChatGPT created clearer test cases and test data, making it better suited and more structured for QA tasks. In contrast.
Gemini performed more like a developer, using strategies typical of non-QA roles. This makes Gemini more useful for QA team members who have a development background.

For integration and usability
ChatGPT is better suited for structured QA tasks such as test planning, requirement analysis, and detailed documentation. 
Gemini, on the other hand, is stronger in AI-assisted debugging and provides automation insights that feel more tailored for developers rather than QA specialists.

In terms of completeness and clarity, 
ChatGPT’s responses were more comprehensive and well-organized, making them easier to use for test documentation. 
Gemini occasionally provided concise answers, which were useful for quick insights, but these sometimes lacked depth.

Overall Review:
ChatGPT is more popular and detail-oriented, making it highly effective for QA professionals. 
Gemini is slightly less popular in comparison but offers more technical insights, which may appeal to users with a development focus.

parwalrahul
Navigator
April 5, 2025

nice ​@NitinMore 

https://testingtitbits.com/
Ensign
April 10, 2025

@parwalrahul  week 2 exercise

Both Chat GPT 4.0 and Gemini are effective and powerful in their ways. However, I found Gemini provides more precise and detailed information when it comes to technical details on applications, tools, and languages.