WOWNEWS24x7

We have integrated a Chatbase chatbot on https://www.wownews24x7.com/. When we ran the same questions through each of the four available models, we observed noticeably different outcomes:

Gemini 2.0

When asking a question, Gemini 2.0 often returned no answer at all. In other words, it didn’t identify any relevant passages and remained blank. Its settings are quite strict—unless a chunk of text matches almost perfectly, Gemini 2.0 won’t use it.

GPT-4

GPT-4 consistently found a relevant excerpt and provided a clear, complete response. Its configuration is more forgiving: it accepts moderately similar passages and examines more pieces of text before answering. As a result, it delivers fully on-topic answers almost every time.

Gemini 1.5

This model typically pulled in a single snippet that shared some overlapping keywords but didn’t capture all the details the question demanded. The answer was partially correct—useful in some cases, but incomplete overall.

Claude 4

Claude 4 behaved much like Gemini 1.5. It surfaced a passage with moderate relevance and then generated an answer that touched on the right subject but left out important information.

Why These Differences

Each model is set up with its own “strictness” level for retrieval and looks at a different number of text chunks. As a result, asking the exact same question can lead to four very different user experiences. In our free Chatbase plan, these are the only models we can choose. If one model (like Gemini 2.0) never finds an answer while another (GPT-4) always does, that inconsistency can confuse users. A blank or partially correct response can also undermine trust and give the impression that we don’t have the information they need.