Search

Tag: #hallucinations34 results

🤖

Local Tool Calling Remains Finicky

User reports unreliable tool calling in local models like Qwen3.5/3.6 27B/35B, Gemma4 26B, GPS-OSS 20B via Open WebUI and LM Studio. Issues include hallucinations, non-existent files, empty outputs, and execution loops despite Unsloth params. Questions if it's model limits or setup errors.

Reddit r/LocalLLaMACommunityApr 18#tool-calling#hallucinations#local-setup
Memvid Hires $800/Day 'AI Bully'

Memvid Hires $800/Day 'AI Bully'

California startup Memvid offers $800 for an 8-hour 'AI bully' role focused on testing leading chatbots' patience and memory. The job entails provoking AI to expose inconsistencies, forgetting, fudging, or hallucinations without meetings or emails. It highlights a unique approach to AI robustness evaluation.

The Guardian TechnologyMediaMar 19#ai-testing#hallucinations#robustness
LLMs Waffle and Err on GOV.UK Queries

LLMs Waffle and Err on GOV.UK Queries

A study of 11 LLMs reveals they rarely refuse GOV.UK government service queries, even when they should, instead providing verbose responses that bury accurate info. When instructed to be concise, chatbots often introduce factual errors. The research questions their trustworthiness for official information.

The Register - AI/MLMediaFeb 19#government-services#hallucinations#verbosity
Page 3 of 4