3.8 KiB
Local LLM tuning (weak / 8k-class models)
Defaults for OpenAI-compatible local servers such as Green Chat Gemma 12B (overloaded-local, ~8k context).
Env
| Variable | Typical local value |
|---|---|
OPENAI_BASE_URL |
http://<host>:8767/v1 |
OPENAI_MODEL |
id from GET /v1/models (e.g. overloaded-local) |
OPENAI_API_KEY |
required non-empty for Enabled() |
PROCESSING_RPM |
60 |
PROCESSING_MAX_RETRIES |
3 |
Restart both api and worker after changing OPENAI_*. Worker logs should say OpenAI enabled base=… model=…, not heuristic completer.
Product enhance model choice
Admin AI role for product enhance should prefer a model that emits message.content (JSON) under budget. Recommended OpenAI: gpt-5.6-luna (high-volume / cost-efficient GPT-5.6 tier).
Reasoning / “code-fast” style models often spend the entire completion budget in reasoning_content, then return finish_reason=length with empty or truncated JSON. Descrybe retries once at MaxTokensEnhanceRetry and may synthesize formula HTML as salvage — that is not live LLM formula compliance.
For GPT-5 / o-series Chat Completions, the client sends max_completion_tokens (not max_tokens), omits custom temperature, and sets reasoning_effort (none for probes, low for long formula-HTML enhance) so visible HTML is not eaten by deep reasoning.
If enhance logs show finish_reason=length with empty/truncated content even after the retry, switch Admin enhance to a lower-reasoning or classic chat model (e.g. gpt-4o-mini).
Hardening defaults (code)
| Constant | Value | Where |
|---|---|---|
DefaultStructuredTemp |
0.2 (capped ≤0.3; omitted on GPT-5/o-series) |
processing/llm_json.go, openai.go |
MaxTokensEnhance |
24576 |
product title + multi-section HTML description JSON |
MaxTokensEnhanceRetry |
40960 |
one-shot bump on finish_reason=length empty/truncated JSON |
MaxTokensSEO |
180 |
meta title/description |
MaxTokensCampaign |
650 |
email JSON |
MaxProductDescRunes |
3500 |
source description in enhance user context |
MaxAttrKeys / MaxAttrValueRunes |
10 / 60 |
attrs in enhance prompt |
MaxBrandInjectRunes |
500 |
brand kit block |
MaxCampaignProducts |
8 |
campaign product list |
Classic OpenAI-compatible chat completions use max_tokens. GPT-5 / o-series use max_completion_tokens (budget includes hidden reasoning tokens). Length-cap handling lives in OpenAIClient.doComplete / CompleteWithOptions.
Prompt style
- Short system prompts with bullet rules + explicit JSON schema
- One few-shot example max
- Brand kit as compact bullets:
Brand:\n- tone: …\n- do: … - Product context: category + name + truncated desc + priority attrs only
JSON reliability
processing.CompleteJSON:
- Call Completer with completion budget + low temperature (classic) / GPT-5.6 param mapping
StripJSONFences/ isolate{…}- On parse fail → one retry with
INVALID. Reply with ONLY one JSON object… - Call sites fall back to originals/templates instead of storing garbage prose as titles
OpenAIClient additionally treats finish_reason=length with empty or unparseable JSON as failure (not success) and retries once at MaxTokensEnhanceRetry.
Ops tips
- Prefer LAN TCP reachability checks (
Test-NetConnection host -Port 8767) over ICMP - Smoke: ai-full-smoke.md, green-chat-llm.md
- CI / no real model: mock-llm.md (
OPENAI_BASE_URL=http://127.0.0.1:18767/v1) — includes translation + processing verification - Locale lists + commented
OPENAI_*/MOCK_LLM_*keys: root.env.example - Do not commit real keys or LAN secrets