Skip to main content
parwalrahul
Navigator
March 13, 2025
Quiz

Week 2 Exercise - Head-to-Head: Evaluating AI Models

  • March 13, 2025
  • 43 replies
  • 854 views

Objective:

Evaluate and compare the ChatGPT 4.0 and the Gemini model on the same task. 

This exercise will help you understand the strengths and limitations of both models.

Steps:

  1. Step 1 – Test using ChatGPT 4.0 (Default Model):
    • Access ChatGPT: Log into AICamp, ChatGPT 4.0 is available by default.
    • Run a Prompt: Use any testing or work-related prompt.
    • Record Results: Document the output, noting aspects like clarity, correctness, and any extra details provided.
  2. Step 2 – Load the Gemini Model:
    • Add Gemini: Navigate to the model integration section on AICamp and add the Gemini (Google) model. (Remember, adding Gemini is free!). Here is the guide to generating a free Gemini API key: Get a Gemini API key  |  Google AI for Developers
      AD_4nXe-0zFGHYb5vJ-BZ_L6OiJeys5rvbjgjTIZpchs4863tqViihie_UkoanvtVElSK8-h3EQ_w4EP6CzsqELhQ4NfOH6SxAJnHhk243-fULAJKowmEN0Ip-eXYKmaXzVRl959lKXIVA?key=snGJktr2mYF7CSmLmKoCWZ8P
    • Verify Integration: Confirm that Gemini has been successfully loaded and is available on your dashboard.
  3. Step 3 – Test Gemini:
    • Run the Same Prompt: Use the identical testing prompt you ran with ChatGPT on the Gemini model.
    • Record Results: Again, document the output focusing on clarity, correctness, and any unique features or differences from ChatGPT.
  4. Step 4 – Compare and Analyze: Create a comparison summary that highlights:
    • Response Quality: What are the differences in how each model responds?
    • Accuracy: Evaluate which output better meets your requirements.
  5. Step 5 – Final Reflection: Summarize your key takeaways in the reply.

43 replies

Ramanan
Ace Pilot
March 13, 2025

Objective:

Evaluate and compare the ChatGPT 4.0 and the Gemini model on the same task. 

This exercise will help you understand the strengths and limitations of both models.

Steps:

  1. Step 1 – Test using ChatGPT 4.0 (Default Model):
    • Access ChatGPT: Log into AICamp, ChatGPT 4.0 is available by default.
    • Run a Prompt: Use any testing or work-related prompt.
    • Record Results: Document the output, noting aspects like clarity, correctness, and any extra details provided.
  2. Step 2 – Load the Gemini Model:
    • Add Gemini: Navigate to the model integration section on AICamp and add the Gemini (Google) model. (Remember, adding Gemini is free!). Here is the guide to generating a free Gemini API key: Get a Gemini API key  |  Google AI for Developers
      AD_4nXe-0zFGHYb5vJ-BZ_L6OiJeys5rvbjgjTIZpchs4863tqViihie_UkoanvtVElSK8-h3EQ_w4EP6CzsqELhQ4NfOH6SxAJnHhk243-fULAJKowmEN0Ip-eXYKmaXzVRl959lKXIVA?key=snGJktr2mYF7CSmLmKoCWZ8P
    • Verify Integration: Confirm that Gemini has been successfully loaded and is available on your dashboard.
  3. Step 3 – Test Gemini:
    • Run the Same Prompt: Use the identical testing prompt you ran with ChatGPT on the Gemini model.
    • Record Results: Again, document the output focusing on clarity, correctness, and any unique features or differences from ChatGPT.
  4. Step 4 – Compare and Analyze: Create a comparison summary that highlights:
    • Response Quality: What are the differences in how each model responds?
    • Accuracy: Evaluate which output better meets your requirements.
  5. Step 5 – Final Reflection: Summarize your key takeaways in the reply.

Hello ​@parwalrahul ,

Both models are strong in their own ways. ChatGPT 4.0 excels at generating well-structured, in-depth responses that feel natural and insightful. Gemini, on the other hand, focuses more on straightforward, factual delivery. If the task requires detailed analysis or a human-like tone, ChatGPT 4.0 is the better choice. If brevity and direct accuracy are the priority, Gemini performs well.

Ultimately, the best model depends on the specific use case.

 

Thanks,

Ramanan 

Hunt the bugs, ensure the hugs. Quality is everything.
Frank Kokoska
Ensign
March 14, 2025

Hello

 

ChatGPT 4.0 and Gemini make a powerful impression but have different strengths. 
ChatGPT 4.0 offers strong text-based capabilities and delivers creative results. Sometimes somewhat biased answers. You have to be able to read and evaluate the answers correctly. Provides answers even when ChatGPT doesn't know - creative.
Gemini has a multimodal design and the integration with Google services provides a versatile and comprehensive user experience. 
I think the choice between the two models depends on the specific user needs, such as the need for multimodal processing, integration with existing tools or the focus on unbiased answers.

 

 

Frank

parwalrahul
Navigator
March 14, 2025

@Ramanan nice takeaway.. my own takeaway is pretty similar to that.

One clear area where gemini excels is image recognition (OCR), especially if we go into native (local) languages… Google really has some edge there.

https://testingtitbits.com/
parwalrahul
Navigator
March 14, 2025

@Frank Kokoska yeah, for people in the google ecosystem or services, Gemini could really stand out.


I feel similar possibilities will become a reality with copilot (backed by ChatGPT). already Microsoft is integrating it with office applications.


Interesting times ahead 🤞

https://testingtitbits.com/
Ensign
March 17, 2025

 

Hello,

So both Chat GPT 4.0 and Gemini are good and powerful in their own ways. However I found Gemini provides more precise and detailed information when it comes to technical details on application, tools, languages.

Can we get the recordings or minutes of meeting for  week2 session?

Bharat2609
Ensign
March 17, 2025

@parwalrahul 

I used a prompt to generate a test strategy document from both ChatGPT4.0 and the Gemini model and obtained some results from it.

ChatGPT 4.0 and Gemini make a powerful impression but have different strengths. 

1.ChatGPT 4.0  provides a more detailed, structured, and technical approach to the test strategy document, making it ideal for in-depth understanding. In contrast, Gemini offers a more narrative and engaging overview, which may be better for stakeholders seeking a broader understanding without delving into technical specifics. 

2.ChatGPT: Presents a well-organized structure with distinct sections and subsections, making it easy to follow and navigate. It offers a methodical layout that covers all key aspects of the testing strategy.Gemini: While organized, it follows a more narrative style. The sections are present but not as clearly separated, resulting in a more fluid but less formal structure.

3.ChatGPT: Uses formal and technical language suited for a professional audience. The tone is clear, direct, and focused on delivering precise information. Gemini: Adopts a conversational and engaging tone, which may appeal to a broader audience but could lack the technical depth expected in formal documentation.

>>>I believe the choice between the two models depends on the user's specific needs, such as the requirement for multimodal processing, compatibility with existing tools, or a focus on providing unbiased answers.

Attachment- testing document

Bharat
parwalrahul
Navigator
March 17, 2025

@KajalS “technical details on application, tools, languages.”

 

what specific differences did you noticed? interested to know more about this.

https://testingtitbits.com/
Ensign
March 17, 2025

@KajalS “technical details on application, tools, languages.”

 

what specific differences did you noticed? interested to know more about this.

Sure Rahul. One of my question was around how can I integrate API tests written in postman with a pipeline on circleCI.

Gemini responded with almost similar steps but along with few examples with detailing like -

  1. .circleci/config.yml  looks like and what each keyword specifies with in the file (for e.g. -e environment.json,  --reporter cli,juni).
  2. Tips like - Important Security Considerations:
  3. Format of environment.json file to store environment variables.

My question was also around Cypress tool, how to start with it.  It replied with basic information like how to install it, key cypress commands, assertions and examples.

Where as Chat GPT provides more theoretical answers(unless specified more accurately) like key features, advantages, etc.

Ensign
March 17, 2025

Both ChatGPT and Gemini are powerful AI tools that provide in-depth insights, though they take different approaches.

  • ChatGPT offers a well-rounded perspective by highlighting free AI tools for accessibility testing, detailing how each tool targets specific functional areas. It also outlines key components of accessibility testing, covering both manual and automated approaches.

  • Gemini focuses on the role of AI in accessibility testing, explaining how AI is integrated into the process. It provides a detailed analysis of the advantages and disadvantages of using AI for accessibility testing.

Conclusion:
Both models deliver comprehensive information, but the choice depends on specific needs. ChatGPT is ideal for those looking for practical tools and a balanced approach, while Gemini is better suited for exploring AI-driven accessibility testing in depth.

parwalrahul
Navigator
March 18, 2025

@Bharat2609 nice try.

Multimodal processing is indeed an amazing possibility and would be the thing that would get more prominence.

We might be finetuning the responses from public LLMs through our custom models and vice versa.

 

thanks for sharing this possibility. Have a nice day!

 

https://testingtitbits.com/