Troubleshooting
A collection of common operational failures and their fixes during local and two-box testing.
Troubleshoot MirrorNeuron
Use this page when a local or two-node runtime fails to start, validate, schedule, or complete work. Start with read-only diagnostics; do not delete Redis state, bundles, or run records as a first response.
Collect evidence before changing state
Run these commands first and save the output with secrets removed:
mn --version
mn runtime health
mn runtime status
mn node listFor a failed workflow, also record the job ID, run ID, timestamp, and exact error text. Use mn job status <job_id> and mn job monitor <job_id> to distinguish a runtime failure from a blueprint requirement or worker failure.
Redis issues
Redis is not running
Symptoms:
- runtime tests fail immediately
mn blueprint run ...hangs or errors
Diagnose:
docker psIf the mirror-neuron-redis container is listed, check it directly:
docker exec mirror-neuron-redis redis-cli pingPONG confirms that the container is responding. If the container is not running, start the managed runtime first:
mn runtime start
mn runtime healthOnly use a manually started Redis container when you intentionally manage Redis outside the MirrorNeuron deployment. Warning: removing a Redis container can destroy the runtime state stored in that container. Back up or preserve the data before any destructive cleanup.
Redis Sentinel two-box smoke says replica did not become online
Symptoms:
remote replica did not become onlineor remote Redis logs show:
Error condition on socket for SYNC: No route to hostCause:
- the remote box cannot route to the local Redis test port
- remote Docker bridge networking cannot reach the local LAN IP
- firewall rules block the test Redis port
Check from the remote box:
nc -vz -w 3 <local-host> 46379If this fails, let the smoke test auto-select the remote side as the initial primary:
python3 mn-system-tests/test_all.py --redis-ha \
--redis-ha-remote-host <remote-host> \
--redis-ha-local-ip <local-host> \
--redis-ha-remote-ip <remote-host>Expected output includes:
Remote cannot reach local Redis at <local-host>:46379; using remote as initial primary.
two_box_post_failover_write_read_okFor direct script control:
cd MirrorNeuron
bash scripts/test_redis_sentinel_two_box_ha.sh \
--remote-host <remote-host> \
--local-ip <local-host> \
--remote-ip <remote-host> \
--remote-network auto \
--initial-primary autoRedis failover returns READONLY or connection errors
During Sentinel promotion, Redis clients can briefly see:
READONLY You can't write against a read only replicaor:
%Redix.ConnectionError{}MirrorNeuron retries reconnectable Redis errors with bounded backoff. If errors persist, check Sentinel:
redis-cli -p 26379 SENTINEL get-master-addr-by-name mirror-neuronExpected output is the current primary host and port.
OpenShell issues
gateway is not reachable
Symptoms:
- transport errors
- connection reset by peer
- jobs fail before worker code starts
Check:
openshell status
openshell sandbox listExpected output includes:
Status: ConnectedReset:
openshell gateway destroy --name openshell
openshell gateway start
openshell statusstale sandboxes slow everything down
Symptoms:
- long provisioning delays
- tiny jobs feel slow
- repeated benchmark runs degrade over time
Clean prime-test sandboxes:
NO_COLOR=1 openshell sandbox list | awk 'NR>1 && index($1, "prime-worker-")==1 {print $1}' | xargs -I{} openshell sandbox delete {}Cluster issues
:nodistribution
Check:
epmd -names
nc -vz 127.0.0.1 4369Fix:
- make sure
epmdis running - use fixed Erlang distribution ports
- verify local firewall rules
Invalid challenge reply
Symptoms:
[error] ** Connection attempt from node :"node2@<remote-host>" rejected. Invalid challenge reply. **- Nodes fail to form a cluster even when IP and ports are fully reachable
Fix:
- This is an Erlang Cookie mismatch. Both nodes must share the exact same secret cookie.
- If you are running nodes on different physical machines, they will auto-generate different cookies by default.
- Set the cookie explicitly on both boxes before starting:
export MN_COOKIE="my_shared_secret"
API port already in use
Symptoms:
Address already in use- the API sidecar fails to start
Fix:
- By default, the FastAPI gateway binds to port
54001. - Stop the conflicting process or choose another API port, for example:
export MN_API_PORT=54002. - Keep the API port distinct from the gRPC port (
55051by default) and the Docker Model Runner/LiteLLM gateway port (4000by default).
runtime node name already in use
Symptoms:
the name mn1@... seems to be in useeaddrinuse
Fix:
- stop the old runtime first
- avoid starting the same box twice
cluster forms but work does not land on both boxes
Possible causes:
- job is too small
- remote bundle sync failed
- stale CLI/control nodes are confusing routing
- one box has less executor capacity
Check:
bash scripts/cluster_cli.sh --box1-ip <box1-host> --box2-ip <box2-host> --self-ip <box1-host> -- inspect nodes
mn node listMonitor issues
monitor shows too many old jobs
This usually means Redis still contains older job metadata.
Options:
- ignore completed jobs with
--running-only - delete old jobs manually if needed
monitor JSON has build noise
Use the CLI command:
mn job monitor <job_id>
If you need all jobs first, run mn job list.
LLM example issues
Gemini API key missing
Symptoms:
- local LLM e2e fails quickly
- cluster LLM harness fails at the first codegen stage
Fix:
export GEMINI_API_KEY="..."Python version mismatch across boxes
Symptoms:
- code works on box 1 but fails on box 2
- typing-related syntax errors on older Python versions
Check:
python3 --versionTry to keep both boxes on a compatible Python version.
When a run feels slower than expected
Common reasons:
- OpenShell provisioning cost
- cold image pulls
- stale gateway state
- large numbers of very tiny executor tasks
- low executor concurrency
If the workflow itself is tiny but runtime is slow, look first at:
- sandbox lifecycle overhead
- gateway health
- whether jobs are being oversharded
Good diagnostic commands
mn node list
mn job status <job_id>
mn job monitor <job_id>
mn job dead-letters <job_id>
openshell status
openshell sandbox list
epmd -names