You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Clicking "Rebuild embedding index" fails, or appears stuck / very slow.
The embedding backend log (e.g. LM Studio, a self-hosted gateway, or any
OpenAI-compatible endpoint) shows what looks like the same document being
sent many times in one batch.
After a failed rebuild, memory_search returns keyword hits only
(vector_hits=0 in the search log), and stays that way across restarts
until a rebuild finally succeeds — i.e. vector search looks permanently
disabled.
Logs contain:
Embedding request failed followed by an HTTP 429 error
(AccountRateLimitExceeded / insufficient_quota etc.)
Root causes — two independent issues (tracked upstream in ReMe)
Rate-limit (429) is never retried on the batch embedding path — [ReMe [Bug]: 'ascii' codec can't encode character '\xe0' in position 137: ordinal not in range(128) #523] LocalEmbeddingStore._call_with_retry only retries TimeoutError / ConnectionError / OSError with exponential backoff. openai.RateLimitError is not a subclass of those, so a 429 falls into the
generic except Exception branch, where only errors with code insufficient_quota are retried (and only when a retry delay is configured).
Any other 429 kills the whole batch instantly.
Because a failed rebuild keeps the needs_reindex gate active (by design, so
the operation can be retried safely), vector search stays disabled until a
rebuild finally succeeds — which looks "permanent" even though the rate
limit itself is transient.
Stale "full-outline" chunks survive upgrades — [ReMe [Feature]: 限制write_file的执行权限 #524]
Older ReMe versions embedded the complete document outline (every heading)
into every chunk. The persisted chunk store
(mem_metadata/file_store/file_chunks_default_v1.jsonl.zst) is versioned by a
static "v1" string that is not tied to the chunker logic, so after upgrading
ReMe, files whose mtime did not change keep their old poisoned chunks. A full
re-embed then sends bloated, near-duplicate content to the embedding API
(the "same document many times" you see in backend logs) and pollutes
retrieval.
Diagnosis (per agent)
Check the gate: in agent.json → running.reme_light_memory_config.needs_reindex. If it is true and no
rebuild is currently running, vector search is disabled.
Check for poisoned chunks ([Feature]: 限制write_file的执行权限 #524): decompress mem_metadata/file_store/file_chunks_default_v1.jsonl.zst and look for files
with an abnormally high chunk count (e.g. a 20 KB file split into 13 chunks,
each repeating the full heading outline). Or re-run the current chunker on
the same file and compare counts — a poisoned store has more chunks than a
fresh chunk.
The 429 is usually an account-level burst limit and is transient. Wait
~5 minutes for cooldown, then re-trigger the rebuild:
UI: Memory → Rebuild embedding index
API: POST /api/agents/{agentId}/memory/reindex?scope=embedding
Each round caches the chunks that succeeded, so every retry embeds fewer
fresh chunks and is less likely to hit the limit again — it converges in 1–2 rounds.
If the provider itself is unhealthy, fix the provider config first, then
re-trigger.
The file watcher re-chunks them incrementally (it keys on mtime). Wait for Applying modified batch ... and Saved N chunks in the log.
Then re-trigger the embedding rebuild once and verify the chunk count dropped
and every chunk now carries a vector.
Verification
needs_reindex is false.
The chunk store contains ~1 chunk per file (e.g. 492 chunks for 378 files)
and 100% of chunks carry embeddings.
A memory search log line shows vector_hits > 0.
Notes
These are workarounds. The actual fixes are tracked at ReMe #523 and ReMe #524 — they affect all
OpenAI-compatible / DashScope / Gemini / Ollama embedding backends, not just
QwenPaw.
When switching embedding models, expect one full re-embed (new vector space);
do it when the account/backend rate limit is quiet.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Symptoms
OpenAI-compatible endpoint) shows what looks like the same document being
sent many times in one batch.
memory_searchreturns keyword hits only(
vector_hits=0in the search log), and stays that way across restartsuntil a rebuild finally succeeds — i.e. vector search looks permanently
disabled.
Embedding request failedfollowed by an HTTP 429 error(
AccountRateLimitExceeded/insufficient_quotaetc.)embedding reindex incomplete: N chunks failedembedding backfill skipped: reason=manual_reindex_requiredRoot causes — two independent issues (tracked upstream in ReMe)
Rate-limit (429) is never retried on the batch embedding path — [ReMe [Bug]: 'ascii' codec can't encode character '\xe0' in position 137: ordinal not in range(128) #523]
LocalEmbeddingStore._call_with_retryonly retriesTimeoutError / ConnectionError / OSErrorwith exponential backoff.openai.RateLimitErroris not a subclass of those, so a 429 falls into thegeneric
except Exceptionbranch, where only errors with codeinsufficient_quotaare retried (and only when a retry delay is configured).Any other 429 kills the whole batch instantly.
Because a failed rebuild keeps the
needs_reindexgate active (by design, sothe operation can be retried safely), vector search stays disabled until a
rebuild finally succeeds — which looks "permanent" even though the rate
limit itself is transient.
Stale "full-outline" chunks survive upgrades — [ReMe [Feature]: 限制write_file的执行权限 #524]
Older ReMe versions embedded the complete document outline (every heading)
into every chunk. The persisted chunk store
(
mem_metadata/file_store/file_chunks_default_v1.jsonl.zst) is versioned by astatic
"v1"string that is not tied to the chunker logic, so after upgradingReMe, files whose mtime did not change keep their old poisoned chunks. A full
re-embed then sends bloated, near-duplicate content to the embedding API
(the "same document many times" you see in backend logs) and pollutes
retrieval.
Diagnosis (per agent)
agent.json→running.reme_light_memory_config.needs_reindex. If it istrueand norebuild is currently running, vector search is disabled.
Embedding request failed+ 429 → rate-limit issue([Bug]: 'ascii' codec can't encode character '\xe0' in position 137: ordinal not in range(128) #523).
embedding reindex incomplete: N chunks failed→ the rebuild failed(usually due to [Bug]: 'ascii' codec can't encode character '\xe0' in position 137: ordinal not in range(128) #523, or an actually unreachable provider).
mem_metadata/file_store/file_chunks_default_v1.jsonl.zstand look for fileswith an abnormally high chunk count (e.g. a 20 KB file split into 13 chunks,
each repeating the full heading outline). Or re-run the current chunker on
the same file and compare counts — a poisoned store has more chunks than a
fresh chunk.
Recovery (workarounds until upstream fixes)
A. Rate-limit / "permanent" disable (#523)
~5 minutes for cooldown, then re-trigger the rebuild:
POST /api/agents/{agentId}/memory/reindex?scope=embeddingfresh chunks and is less likely to hit the limit again — it converges in
1–2 rounds.
re-trigger.
B. Poisoned chunks (#524)
watched markdown files (content unchanged):
Applying modified batch ...andSaved N chunksin the log.and every chunk now carries a vector.
Verification
needs_reindexisfalse.and 100% of chunks carry embeddings.
vector_hits > 0.Notes
ReMe #523 and
ReMe #524 — they affect all
OpenAI-compatible / DashScope / Gemini / Ollama embedding backends, not just
QwenPaw.
do it when the account/backend rate limit is quiet.
All reactions