Skip to main content
parwalrahul
Navigator
March 13, 2025
Quiz

Week 2 Exercise - Head-to-Head: Evaluating AI Models

  • March 13, 2025
  • 43 replies
  • 854 views

Objective:

Evaluate and compare the ChatGPT 4.0 and the Gemini model on the same task. 

This exercise will help you understand the strengths and limitations of both models.

Steps:

  1. Step 1 – Test using ChatGPT 4.0 (Default Model):
    • Access ChatGPT: Log into AICamp, ChatGPT 4.0 is available by default.
    • Run a Prompt: Use any testing or work-related prompt.
    • Record Results: Document the output, noting aspects like clarity, correctness, and any extra details provided.
  2. Step 2 – Load the Gemini Model:
    • Add Gemini: Navigate to the model integration section on AICamp and add the Gemini (Google) model. (Remember, adding Gemini is free!). Here is the guide to generating a free Gemini API key: Get a Gemini API key  |  Google AI for Developers
      AD_4nXe-0zFGHYb5vJ-BZ_L6OiJeys5rvbjgjTIZpchs4863tqViihie_UkoanvtVElSK8-h3EQ_w4EP6CzsqELhQ4NfOH6SxAJnHhk243-fULAJKowmEN0Ip-eXYKmaXzVRl959lKXIVA?key=snGJktr2mYF7CSmLmKoCWZ8P
    • Verify Integration: Confirm that Gemini has been successfully loaded and is available on your dashboard.
  3. Step 3 – Test Gemini:
    • Run the Same Prompt: Use the identical testing prompt you ran with ChatGPT on the Gemini model.
    • Record Results: Again, document the output focusing on clarity, correctness, and any unique features or differences from ChatGPT.
  4. Step 4 – Compare and Analyze: Create a comparison summary that highlights:
    • Response Quality: What are the differences in how each model responds?
    • Accuracy: Evaluate which output better meets your requirements.
  5. Step 5 – Final Reflection: Summarize your key takeaways in the reply.

43 replies

parwalrahul
Navigator
March 18, 2025

@KajalS nice observation. thanks for the details.

even I have a feeling that gemini is relatively less used compared to the quality it offers.

it is at part with GPT models and sometimes even better in specific tech tasks.

 

Let’s see where this race of models will end. as of now, there is no clear winner.

https://testingtitbits.com/
parwalrahul
Navigator
March 18, 2025

thanks for sharing your response and observations, ​@Darshana :)

https://testingtitbits.com/
Ensign
March 18, 2025
Feature GPT-4-mini Gemini Flash 1.5
Structure Highly structured, clear sections Less structured, more narrative style
Investigative Depth Less speculative, focuses on facts Speculates on possible causes
Overall Quality Excellent, ready for immediate action Good, but requires further investigation

 

GPT-4-mini provides a superior & clear structure

Gemini Flash 1.5 offers valuable context but lacks the precision

 

parwalrahul
Navigator
March 18, 2025

nice comparison and to the point. great job, ​@satishracherla 

https://testingtitbits.com/
Ensign
March 19, 2025

Hello ​@parwalrahul 

 

I have tried to use two examples in the attached document to create a user story with identical prompt across both the LLM's Chat GPT and Gemini

Comparison Summary:

Response Quality:

  1. Structure:

    • ChatGPT: Provides a more traditional Jira ticket structure with clear sections.
    • Gemini: Offers a more detailed and comprehensive Jira ticket format.
  2. Clarity:

    • ChatGPT: Concise and straightforward, but lacks some detail.
    • Gemini: More thorough and descriptive, providing clearer context.
  3. Detail Level:

    • ChatGPT: Offers basic information, somewhat generic.
    • Gemini: Provides more specific details and scenarios.
  4. BDD Format:

    • ChatGPT: Does not strictly adhere to BDD format in acceptance criteria.
    • Gemini: Correctly uses Given-When-Then format for scenarios.

Accuracy:

  1. User Story:

    • ChatGPT: Presents a basic user story format.
    • Gemini: Provides a more comprehensive user story with clear benefit.
  2. Acceptance Criteria:

    • ChatGPT: Lists criteria but doesn't follow BDD format.
    • Gemini: Correctly uses BDD format with specific scenarios.
  3. Definition of Ready (DoR):

    • ChatGPT: Includes relevant points but is somewhat generic.
    • Gemini: More specific to the task, includes technical considerations.
  4. Definition of Done (DoD):

    • ChatGPT: Covers essential points but lacks some technical specifics.
    • Gemini: More comprehensive, includes testing and deployment steps.
  5. Additional Elements:

    • ChatGPT: Includes priority and labels.
    • Gemini: Adds project name, issue type, assignee, and epic link.

Unique Features:

  • ChatGPT: Includes a section for user feedback mechanisms.
  • Gemini: Provides negative case scenarios and more technical details in DoD.

Overall Evaluation:

While both outputs provide valuable information, Gemini's response appears to better meet the requirements of the prompt. It offers a more structured, detailed, and technically accurate representation of a Jira ticket for a user story. The use of proper BDD format in the acceptance criteria and the inclusion of both positive and negative scenarios demonstrate a more comprehensive understanding of the task.

ChatGPT's response, while clear and concise, lacks some of the technical depth and specificity that would be expected in a real-world Jira ticket for this type of task.

In terms of accuracy and adherence to best practices in Agile development and Jira ticket creation, Gemini's output is superior in this instance.

parwalrahul
Navigator
March 19, 2025

@ameet213 true. even i have felt this.

your observations also matches the observations by a lot of other members who attended this exercise.

Also, with gemini, I have noticed that it has some guards or mechanism to stop answering if user puts in excessive debug logs or stack trace and asks for more information.

this was an interesting observation that I had. rest, it’s similar to yours. thanks :)

https://testingtitbits.com/
Specialist
March 19, 2025

Hi Rahul,

 

Couple of things I want to highlight:

AICamp is definitely worth the try and its going to be the camp I can refer to check different models or to save the prompts from different models at a place. 

  1. Utilized GPT 4o-mini and Gemini - Both are good.
  2. Created an assistant but it can do better still there is a lot of improvement and it’s from the instructions from user side and also the template which assistant is following.
  3. We can share the workspace with any one - this feature people won’t expect.
  4. These models were not that powerful when compared to accessing them in their own websites like I asked it to generate an image but it doesn’t(both models) - Text generation is what these models are for. And we can also chat with a document.

Here are my comparisons with an assistant:

  1. Clarity: Created some test cases by using GPT 4o(openAI), 1.5 pro(Gemini) but the response provided from Gemini is good when compared to GPT 4o
  2. Correctness: Can’t comment on this both are not connected to the internet as I can see their trained data is limited to OCT 2023 but till this point, everything is ok.
  3. Consistency: In multiple test runs with the same prompt, Gemini provided more consistent and predictable responses compared to GPT 4o-mini. GPT 4o-mini exhibited greater variability in its outputs, sometimes deviating significantly from previous responses.
  4. Image generated from Imagen3 by Google

     

Found one thing interesting while testing it: That's great! I'm an AI assistant, created by AICamp.

So I believe AICamp is having it’s own template and using the API’s from OpenAI and Google.

 

If you find the pros of unified access and prompt management within AICamp more interesting, it's the platform for you. If the limitations (cons) outweigh the benefits, then the respective GenAI websites are the best to use.

Kusumketu
Ensign
March 20, 2025

Objective:

Evaluate and compare the ChatGPT 4.0 and the Gemini model on the same task. 

This exercise will help you understand the strengths and limitations of both models.

Steps:

  1. Step 1 – Test using ChatGPT 4.0 (Default Model):
    • Access ChatGPT: Log into AICamp, ChatGPT 4.0 is available by default.
    • Run a Prompt: Use any testing or work-related prompt.
    • Record Results: Document the output, noting aspects like clarity, correctness, and any extra details provided.
  2. Step 2 – Load the Gemini Model:
    • Add Gemini: Navigate to the model integration section on AICamp and add the Gemini (Google) model. (Remember, adding Gemini is free!). Here is the guide to generating a free Gemini API key: Get a Gemini API key  |  Google AI for Developers
      AD_4nXe-0zFGHYb5vJ-BZ_L6OiJeys5rvbjgjTIZpchs4863tqViihie_UkoanvtVElSK8-h3EQ_w4EP6CzsqELhQ4NfOH6SxAJnHhk243-fULAJKowmEN0Ip-eXYKmaXzVRl959lKXIVA?key=snGJktr2mYF7CSmLmKoCWZ8P
    • Verify Integration: Confirm that Gemini has been successfully loaded and is available on your dashboard.
  3. Step 3 – Test Gemini:
    • Run the Same Prompt: Use the identical testing prompt you ran with ChatGPT on the Gemini model.
    • Record Results: Again, document the output focusing on clarity, correctness, and any unique features or differences from ChatGPT.
  4. Step 4 – Compare and Analyze: Create a comparison summary that highlights:
    • Response Quality: What are the differences in how each model responds?
    • Accuracy: Evaluate which output better meets your requirements.
  5. Step 5 – Final Reflection: Summarize your key takeaways in the reply.

Here we go on Week 2 assignment:

 

Comparison and Analysis: I gave Prompt to write Test Case for Air Ticket Booking System

GPT-4o-mini: Detailed Structured Test Case

Strengths:

Provides a clear and detailed structure using a step-by-step approach.

Includes well-defined sections like Test Case ID, Description, Preconditions, Test Data, Steps, and Expected Results.

Ensures traceability and ease of execution for testers.

Suitable for manual testing and easy to maintain.

Limitations:

Limited coverage compared to Gemini; it focuses on one scenario.

Lacks broader test coverage including negative and security test cases.

Not scalable for complex systems with multiple scenarios.
 

Gemini: Categorized Test Cases

Strengths:

Offers a comprehensive test coverage across functional, negative, usability, and security aspects.

Organized by categories, making it easy to prioritize and manage testing efforts.

Ideal for large-scale applications where multiple components are involved.

Encourages collaboration with QA, Dev, and Product teams.

Limitations:

Test steps are not elaborated, which may require additional context for testers.

Requires separate detailed steps and data for each test case.

May introduce inconsistencies if not standardized.

 

Response Quality Comparison between GPT-4o-mini and Gemini

GPT-4o-mini provides a clear and detailed test scenario with actionable steps. It’s ideal for step-by-step validation.

Gemini offers a holistic view, ensuring better test coverage. It’s suitable for teams aiming to cover edge cases and ensure overall system stability.

Accuracy Analysis

GPT-4o-mini is more accurate for functional validation of a specific scenario.

Gemini is more accurate for identifying defects in various parts of the system, including performance, usability, and security.

Final Reflection

For simple scenarios or when precision is needed, GPT-4o-mini is more appropriate.

For complex systems requiring end-to-end coverage, Gemini is recommended as it ensures a broader and more thorough testing approach.

A hybrid approach, using GPT-4o-mini for critical test scenarios and Gemini for comprehensive coverage, would offer the most effective results.

** Attached testcase docs for reference

Kusum Ketu
Ensign
March 20, 2025

My prompt to ChatGPT and Gemini was, "Teach me Playwright." I noticed that Gemini focused more on the theoretical aspects, while ChatGPT provided results that were more technically oriented. One interesting thing I observed about ChatGPT, which I didn't see in Gemini, was that ChatGPT's responses included prompts for me to answer at the end. It felt like it was trying to interact with you.

parwalrahul
Navigator
March 20, 2025

@Dinesh_Gujarathi thanks for your detailed submission brother :)


you really did cover the exercise from all the aspects including setting up stuff via AICamp.

Because you tried that, I can see how you could instantly realize it’s potential and amazing capabilities.


I use it regularly and I know of the image generation bug too. they are working on it :D (bug reported by me already).


I have also setup a couple of assistants via it and I am excited to expand them more as the knowledge feature for it gets enabled.

 

See you in the event today. have a nice evening!

https://testingtitbits.com/