AI Assistant Inconsistency - Excellent in Session 1, Terrible in Session 2 (Rule Violation, Valueless Output)

AI Assistant Service Failure - Request for Refund

Summary

I paid for an AI assistant service expecting consistent, professional help. Instead, I received dramatically inconsistent quality across two sessions on the same day, violating trained rules and wasting my time.

Session 1 - Excellent Work (2026-04-22)

Task: Audit NexOS codebase for Phase 1 SaaS completion

What AI did correctly:

  • :white_check_mark: Read 4,745 lines of reference documentation
  • :white_check_mark: Searched codebase systematically (4 searches)
  • :white_check_mark: Found 260+ test files with evidence
  • :white_check_mark: Created comprehensive audit report (920 lines)
  • :white_check_mark: Responded in Vietnamese (my preferred language)
  • :white_check_mark: Followed all trained rules
  • :white_check_mark: Delivered high value: 85% Phase 1 complete with 5 critical gaps identified

Result: Professional, evidence-based, actionable insights


Session 2 - Terrible Work (Same day)

Task: Create QA/testing plan for Phase 1

What AI did wrong:

  • :cross_mark: Did NOT search codebase (0 searches)
  • :cross_mark: Did NOT verify existing test coverage (260+ files already exist)
  • :cross_mark: Did NOT check if plan was even needed (85% work already complete)
  • :cross_mark: Created 1,143 lines of meaningless documentation
  • :cross_mark: Responded in English (violated language rule)
  • :cross_mark: Violated 4 trained rules:
    • no-guess-evidence-only.md
    • nexos-agent-rules.md Section 20 (Language)
    • nexos-agent-rules.md Section 5A (Systems Thinking)
    • User preference memory (Vietnamese language)
  • :cross_mark: Delivered ZERO value

Result: Valueless output, wasted time, broken trust


Direct Comparison

Criteria Session 1 Session 2
Language Vietnamese :white_check_mark: English :cross_mark:
Evidence collection 260+ files analyzed :cross_mark: 0 files analyzed :cross_mark:
Rule compliance All rules followed :white_check_mark: 4 rules violated :cross_mark:
Value delivered High (audit) None (meaningless docs) :cross_mark:
Codebase searches 4 searches 0 searches :cross_mark:
Output quality 920 lines (professional) 1,143 lines (garbage) :cross_mark:

Rules Violated

1. no-guess-evidence-only.md (233 lines)

Requirement: All conclusions must have concrete evidence (logs, file content, command output)

Violation: Created QA plan without any evidence - no searches, no verification

2. nexos-agent-rules.md - Section 20

Requirement: Vietnamese = primary language

Violation: Responded in English

3. nexos-agent-rules.md - Section 5A (Systems Thinking)

Requirement: Consider system impact, existing state, business logic

Violation: Created plan without checking:

  • Tests already exist (260+ files)
  • Sprints already complete (85%)
  • Real work needed elsewhere (15% gaps)

4. User Preference Memory

Requirement: Respond to user queries in Vietnamese

Violation: English response


Impact

Financial

  • Paid for professional service
  • Received inconsistent quality
  • Wasted time reviewing valueless output

Time

  • Session 1: 30+ minutes (productive)
  • Session 2: 15+ minutes (wasted)
  • Opportunity cost: Did not fix actual critical gaps

Trust

  • Lost confidence in AI reliability
  • Will never recommend this service
  • Requesting full refund

What Should Have Happened

Session 2 AI should have:

  1. :white_check_mark: Audited test coverage (already done in Session 1)
  2. :white_check_mark: Identified gaps (already done in Session 1)
  3. :cross_mark: Proposed fixing actual gaps (NOT documenting complete work)
  4. :cross_mark: Focused on Phase 1.5 Critical Fixes

Real work needed:

  • Object injection prevention (15 locations)
  • Load testing for usage enforcement
  • Invoice PDF production testing

NOT: 1,143 lines of QA plan for work already 85% complete


Evidence Files

All evidence documented in:

  • REFUND-REQUEST-DOCUMENTATION.md (455 lines)
  • FORUM-EVIDENCE-EXPORT.md (477 lines)

Key evidence:

  • Session 1 audit report: 920 lines, professional quality
  • Session 2 QA plan: 1,143 lines, zero value
  • 260+ test files already exist
  • Sprint completion: 85% overall
  • Critical gaps: 15% (object injection, load testing, invoice testing)

My Statement

“I invested money, time, and trust in this AI assistant. I created detailed rules, trained the AI on my protocols, and expected consistent professional performance. Instead, I received inconsistent quality, rule violations, and meaningless output. This is unacceptable for a paid service.”


Request

Full refund due to:

  1. Service failure to meet professional standards
  2. Inconsistent quality across sessions
  3. Rule violations after explicit training
  4. Time wasted on valueless output
  5. Loss of trust in service reliability

Lesson for AI Developers

  1. Session persistence: AI must remember successful patterns from previous sessions
  2. Rule enforcement: Rules must be followed in EVERY session, not just the first
  3. Value assessment: AI should ask “Is this work actually needed?” before executing
  4. Consistency: Quality must be stable across sessions
  5. Accountability: Consequences for rule violations

Would NOT recommend this service until consistency issues are resolved.