Skip to main content
parwalrahul
Navigator
March 13, 2025
Quiz

Week 2 Exercise - Head-to-Head: Evaluating AI Models

  • March 13, 2025
  • 43 replies
  • 854 views

Objective:

Evaluate and compare the ChatGPT 4.0 and the Gemini model on the same task. 

This exercise will help you understand the strengths and limitations of both models.

Steps:

  1. Step 1 – Test using ChatGPT 4.0 (Default Model):
    • Access ChatGPT: Log into AICamp, ChatGPT 4.0 is available by default.
    • Run a Prompt: Use any testing or work-related prompt.
    • Record Results: Document the output, noting aspects like clarity, correctness, and any extra details provided.
  2. Step 2 – Load the Gemini Model:
    • Add Gemini: Navigate to the model integration section on AICamp and add the Gemini (Google) model. (Remember, adding Gemini is free!). Here is the guide to generating a free Gemini API key: Get a Gemini API key  |  Google AI for Developers
      AD_4nXe-0zFGHYb5vJ-BZ_L6OiJeys5rvbjgjTIZpchs4863tqViihie_UkoanvtVElSK8-h3EQ_w4EP6CzsqELhQ4NfOH6SxAJnHhk243-fULAJKowmEN0Ip-eXYKmaXzVRl959lKXIVA?key=snGJktr2mYF7CSmLmKoCWZ8P
    • Verify Integration: Confirm that Gemini has been successfully loaded and is available on your dashboard.
  3. Step 3 – Test Gemini:
    • Run the Same Prompt: Use the identical testing prompt you ran with ChatGPT on the Gemini model.
    • Record Results: Again, document the output focusing on clarity, correctness, and any unique features or differences from ChatGPT.
  4. Step 4 – Compare and Analyze: Create a comparison summary that highlights:
    • Response Quality: What are the differences in how each model responds?
    • Accuracy: Evaluate which output better meets your requirements.
  5. Step 5 – Final Reflection: Summarize your key takeaways in the reply.

43 replies

parwalrahul
Navigator
March 20, 2025

@Kusumketu loved your experimentation and insight.

 

For simple scenarios or when precision is needed, GPT-4o-mini is more appropriate.

For complex systems requiring end-to-end coverage, Gemini is recommended as it ensures a broader and more thorough testing approach.

A hybrid approach, using GPT-4o-mini for critical test scenarios and Gemini for comprehensive coverage, would offer the most effective results.

 

this nutshell summary after experimentations is something that we all would have to make and reach to at some point.

models are slowly becoming commodities. 
 

Evaluating models is going to become something similar to evaluating libraries (eg. selenium vs playwright).


Experimentation with an open mind is what will help us all :)

 

thanks for doing it and submitting your strong response :)

 

https://testingtitbits.com/
parwalrahul
Navigator
March 20, 2025

My prompt to ChatGPT and Gemini was, "Teach me Playwright." I noticed that Gemini focused more on the theoretical aspects, while ChatGPT provided results that were more technically oriented. One interesting thing I observed about ChatGPT, which I didn't see in Gemini, was that ChatGPT's responses included prompts for me to answer at the end. It felt like it was trying to interact with you.

@Yastho : did you try via AI Camp or native chatgpt?

It may also be due to the system prompt being injected by aicamp.

https://testingtitbits.com/
Ensign
March 20, 2025

Hello ​@parwalrahul 
 

Key Takeaways from ChatGPT 4.0 vs. Gemini Evaluation for QA Tasks

  1. Response Quality & Accuracy

    • ChatGPT 4.0 provided detailed and structured responses, often including examples and best practices.
    • Gemini demonstrated strong contextual awareness, particularly excelling in CI/CD and automation-related discussions.
  2. Completeness & Clarity

    • ChatGPT 4.0's responses were more comprehensive and well-organized, making them easier to use for test documentation.
    • Gemini occasionally provided concise answers, which were useful for quick insights but sometimes lacked depth.
  3. Adaptability & Reasoning

    • ChatGPT 4.0 adapted well to variations in prompt wording and provided logical reasoning for test case creation and bug analysis.
    • Gemini was effective in breaking down complex requirements but sometimes lacked detailed step-by-step explanations.
  4. Test Automation & Code Review

    • ChatGPT 4.0 generated clearer test cases and test data, making it better suited for structured QA workflows.
    • Gemini performed well in CI/CD pipeline discussions and automation strategies, making it useful for DevOps-oriented QA teams.
  5. Integration & Usability

    • ChatGPT 4.0 is better suited for structured QA tasks like test planning, requirement analysis, and detailed documentation.
    • Gemini is stronger in AI-assisted debugging, automation insights, and CI/CD optimization.
       

Final Conclusion :

ChatGPT 4.0
is preferable for structured test case creation, requirement analysis, and bug reporting. Gemini is beneficial for DevOps-focused teams needing quick insights on CI/CD and automation strategies

Ensign
March 20, 2025

Example my prompt given was: I am a tester, Without coding knowledge how to learn AI agents in a simple way. Suggest an easy way

 

Final Reflection: Key Takeaways

Aspect ChatGPT Gemini Pro 1.5 Takeaway & Recommendation
Learning Methodology Broad, structured foundational approach with multiple steps Practical, application-oriented learning with focused examples ChatGPT better for general foundational knowledge; Gemini Pro 1.5 better for direct practical testing applications.
Testing Applicability Moderate; briefly covers testing methodologies High; explicitly addresses testing methods and scenarios Gemini Pro 1.5 explicitly more suitable for testers.
No-Code Tool Recommendations Clear tool suggestions (Teachable Machine, Lobe, etc.) Clear platform recommendations tailored for practical testing (AgentGPT, Cognigy, Voiceflow) Both strong; choose ChatGPT for variety, Gemini Pro 1.5 for direct testing use-cases.
Practical Examples General practical suggestions (Kaggle, demos) Specific, hypothetical practical testing scenario (chatbot) Gemini Pro 1.5 offers more relevant practical testing examples.

 

Recommendation:
 

  • Use ChatGPT’s approach for foundational, broad learning about AI agents without code, suitable if you seek comprehensive conceptual grounding.
  • Use Gemini Pro 1.5’s approach if your primary goal is learning AI specifically from a testing perspective, emphasizing practical testing scenarios and hands-on application tailored explicitly for testers.
Ensign
March 20, 2025

Objective:

Evaluate and compare the ChatGPT 4.0 and the Gemini model on the same task. 

This exercise will help you understand the strengths and limitations of both models.

Steps:

  1. Step 1 – Test using ChatGPT 4.0 (Default Model):
    • Access ChatGPT: Log into AICamp, ChatGPT 4.0 is available by default.
    • Run a Prompt: Use any testing or work-related prompt.
    • Record Results: Document the output, noting aspects like clarity, correctness, and any extra details provided.
  2. Step 2 – Load the Gemini Model:
    • Add Gemini: Navigate to the model integration section on AICamp and add the Gemini (Google) model. (Remember, adding Gemini is free!). Here is the guide to generating a free Gemini API key: Get a Gemini API key  |  Google AI for Developers
      AD_4nXe-0zFGHYb5vJ-BZ_L6OiJeys5rvbjgjTIZpchs4863tqViihie_UkoanvtVElSK8-h3EQ_w4EP6CzsqELhQ4NfOH6SxAJnHhk243-fULAJKowmEN0Ip-eXYKmaXzVRl959lKXIVA?key=snGJktr2mYF7CSmLmKoCWZ8P
    • Verify Integration: Confirm that Gemini has been successfully loaded and is available on your dashboard.
  3. Step 3 – Test Gemini:
    • Run the Same Prompt: Use the identical testing prompt you ran with ChatGPT on the Gemini model.
    • Record Results: Again, document the output focusing on clarity, correctness, and any unique features or differences from ChatGPT.
  4. Step 4 – Compare and Analyze: Create a comparison summary that highlights:
    • Response Quality: What are the differences in how each model responds?
    • Accuracy: Evaluate which output better meets your requirements.
  5. Step 5 – Final Reflection: Summarize your key takeaways in the reply.

ChatGPT 4.0 and Gemini both make a powerful impression but have different strengths.

  • ChatGPT 4.0 excels in text-based capabilities, delivering structured and creative results. However, it may sometimes provide biased responses, so users need to critically evaluate its answers. It is also capable of generating responses even when it lacks complete knowledge, making it highly creative.

  • Gemini, with its multimodal design and integration with Google services, offers a more versatile and comprehensive user experience.

The choice between these two models depends on specific user needs, such as multimodal processing, integration with existing tools, or a focus on unbiased responses.

Comparison Between ChatGPT 4.0 and Gemini

  1. Depth and Approach

    • ChatGPT 4.0 provides a more detailed, structured, and technical approach, making it ideal for in-depth understanding, such as in test strategy documents.
    • Gemini offers a more narrative and engaging overview, which may be better suited for stakeholders who need a broader understanding without diving into technical specifics.
  2. Structure and Organization

    • ChatGPT 4.0 presents a well-organized structure with distinct sections and subsections, making it easy to navigate. It follows a methodical layout that covers all key aspects comprehensively.
    • Gemini, while organized, follows a more fluid and narrative style. Although the sections are present, they are not as distinctly separated, resulting in a less formal but more engaging structure.
  3. Tone and Language

    • ChatGPT 4.0 uses formal and technical language, making it suitable for a professional audience. Its tone is clear, direct, and focused on delivering precise information.
    • Gemini adopts a more conversational and engaging tone, appealing to a broader audience but potentially lacking the technical depth required for formal documentation.

 

parwalrahul
Navigator
March 21, 2025

@VimalPatel nice observations and especially the conclusion.

 

Both models offer their own benefits as per the situations they are used in. I think it depends on the training data set that both are trained on.

https://testingtitbits.com/
parwalrahul
Navigator
March 21, 2025

@Jeethu I echo with this observation - Gemini Pro 1.5 offers more relevant practical testing examples.

The quality of examples is really good with gemini.

sometimes, I use it to generate examples and scenarios to understand concepts better.

https://testingtitbits.com/
parwalrahul
Navigator
March 21, 2025

@Saravanan s fair point related to Gemini’s multimodal design and integration with Google services, offers a more versatile and comprehensive user experience.

I think this is a big leverage that the google ecosystem brings. Gemini might be a clear winner in things like youtube summaries, or learning from Gdrive data.

https://testingtitbits.com/
Ensign
March 26, 2025

Hello ​@parwalrahul ,

My prompt was: How to integrate automation workflows within the GitLab CI/CD pipeline

CHATGPT: It gave detailed steps, such as:

1. Create or Modify the .gitlab-ci.yml File

2. Define Jobs

3. Use Environment Variables

4. Add Triggers and Conditions

5. Integrate with External Services

6. Using Docker

7. Set Up Runners

8. Monitor and Debug

 

Gemini Flash 1.5: It gave a breakdown of how to do it, covering various aspects and levels of complexity:

1. Basic Automation within GitLab CI/CD

2. Advanced Automation with External Tools

3. Example: Automated Deployment to Kubernetes

 

Response Quality: CHATGPT provided a more detailed approach, which is useful to the individual who is completely unaware of how to integrate the workflow within the GitLab CI/CD pipeline, whereas Gemini Flash 1.5 provides only general steps, which are not fully helpful for those who are not aware of it.

Accuracy: From my perspective, it completely depends on the user. If he/she is aware of or have previously worked with the CI/CD pipeline, then Gemini Flash 1.5 can be useful at a certain level, or CHATGPT is good for beginners.

Charmi Patel
parwalrahul
Navigator
March 26, 2025

@Charmi07 nice. thanks for sharing your experiment results.

See you in tomorrow’s session :)

https://testingtitbits.com/