Files
descrybe/docs/local-llm-tuning.md
2026-08-23 12:40:59 +02:00

3.8 KiB

Local LLM tuning (weak / 8k-class models)

Defaults for OpenAI-compatible local servers such as Green Chat Gemma 12B (overloaded-local, ~8k context).

Env

Variable Typical local value
OPENAI_BASE_URL http://<host>:8767/v1
OPENAI_MODEL id from GET /v1/models (e.g. overloaded-local)
OPENAI_API_KEY required non-empty for Enabled()
PROCESSING_RPM 60
PROCESSING_MAX_RETRIES 3

Restart both api and worker after changing OPENAI_*. Worker logs should say OpenAI enabled base=… model=…, not heuristic completer.

Product enhance model choice

Admin AI role for product enhance should prefer a model that emits message.content (JSON) under budget. Recommended OpenAI: gpt-5.6-luna (high-volume / cost-efficient GPT-5.6 tier).

Reasoning / “code-fast” style models often spend the entire completion budget in reasoning_content, then return finish_reason=length with empty or truncated JSON. Descrybe retries once at MaxTokensEnhanceRetry and may synthesize formula HTML as salvage — that is not live LLM formula compliance.

For GPT-5 / o-series Chat Completions, the client sends max_completion_tokens (not max_tokens), omits custom temperature, and sets reasoning_effort (none for probes, low for long formula-HTML enhance) so visible HTML is not eaten by deep reasoning.

If enhance logs show finish_reason=length with empty/truncated content even after the retry, switch Admin enhance to a lower-reasoning or classic chat model (e.g. gpt-4o-mini).

Hardening defaults (code)

Constant Value Where
DefaultStructuredTemp 0.2 (capped ≤0.3; omitted on GPT-5/o-series) processing/llm_json.go, openai.go
MaxTokensEnhance 24576 product title + multi-section HTML description JSON
MaxTokensEnhanceRetry 40960 one-shot bump on finish_reason=length empty/truncated JSON
MaxTokensSEO 180 meta title/description
MaxTokensCampaign 650 email JSON
MaxProductDescRunes 3500 source description in enhance user context
MaxAttrKeys / MaxAttrValueRunes 10 / 60 attrs in enhance prompt
MaxBrandInjectRunes 500 brand kit block
MaxCampaignProducts 8 campaign product list

Classic OpenAI-compatible chat completions use max_tokens. GPT-5 / o-series use max_completion_tokens (budget includes hidden reasoning tokens). Length-cap handling lives in OpenAIClient.doComplete / CompleteWithOptions.

Prompt style

  • Short system prompts with bullet rules + explicit JSON schema
  • One few-shot example max
  • Brand kit as compact bullets: Brand:\n- tone: …\n- do: …
  • Product context: category + name + truncated desc + priority attrs only

JSON reliability

processing.CompleteJSON:

  1. Call Completer with completion budget + low temperature (classic) / GPT-5.6 param mapping
  2. StripJSONFences / isolate {…}
  3. On parse fail → one retry with INVALID. Reply with ONLY one JSON object…
  4. Call sites fall back to originals/templates instead of storing garbage prose as titles

OpenAIClient additionally treats finish_reason=length with empty or unparseable JSON as failure (not success) and retries once at MaxTokensEnhanceRetry.

Ops tips

  • Prefer LAN TCP reachability checks (Test-NetConnection host -Port 8767) over ICMP
  • Smoke: ai-full-smoke.md, green-chat-llm.md
  • CI / no real model: mock-llm.md (OPENAI_BASE_URL=http://127.0.0.1:18767/v1) — includes translation + processing verification
  • Locale lists + commented OPENAI_* / MOCK_LLM_* keys: root .env.example
  • Do not commit real keys or LAN secrets