Skip to content

Connecting to a cluster that requires 2FA

andy edited this page Sep 25, 2026 · 1 revision

Connecting to a cluster that requires two-factor authentication

ECCE submits jobs by running ssh to the machine you registered. Most computing centres now require two-factor authentication, and ECCE has no way to answer a 2FA prompt itself.

You do not have to give up job submission because of this. If you authenticate before starting ECCE and leave that authenticated connection open, ECCE can use it, and everything works normally — submission, live monitoring and output retrieval.

Two ways to do it. Try connection sharing first; it is simpler and needs nothing running in the background.


Option 1 — connection sharing (recommended)

OpenSSH can reuse one authenticated connection for every later ssh to the same host. You do the 2FA once, in your own terminal, and ECCE's connections ride on it without prompting.

Put this in ~/.ssh/config:

Host myhpc
    HostName login.hpc.example.edu
    User your-cluster-username
    ControlMaster auto
    ControlPath ~/.ssh/cm-%r@%h:%p
    ControlPersist 8h

ControlPersist 8h keeps the shared connection alive for eight hours after the last use — long enough for a working day. Adjust to taste.

Then, before you start ECCE:

ssh myhpc

Complete the 2FA as usual. You can close that terminal window; the shared connection stays up.

Register the machine in ECCE as myhpc — the alias, not the real host name — and give the same username you put in the config.

To check the shared connection is alive:

ssh -O check myhpc

To drop it deliberately:

ssh -O exit myhpc

If your centre disallows it. Some sites set ControlMaster no in the system-wide configuration, or require 2FA on every channel. In that case connection sharing will re-prompt and you want option 2.


Option 2 — an SSH tunnel

Open a tunnel yourself, authenticate through it, and point ECCE at the local end.

ssh -f -N -L 2222:login.hpc.example.edu:22 gateway.hpc.example.edu

Then add an alias for the local end:

Host hpc-tunnel
    HostName 127.0.0.1
    Port 2222
    User your-cluster-username

Register the machine in ECCE as hpc-tunnel.

Do not register the machine as localhost

This matters, because getting it wrong fails quietly and can give you results from the wrong computer.

ECCE decides whether a machine is remote by looking at its name. localhost, localhost.localdomain, 127.0.0.1 and ::1 all mean this machine, and for those it does not use ssh at all — it runs the job locally. Pointing ECCE at a forwarded port on localhost can therefore run your calculation on your own workstation instead of on the cluster.

Worse, whether that happens depends on your username:

  • if the username you registered differs from your local one, ECCE treats the machine as remote and the tunnel is used, so it works;
  • if they match, ECCE runs everything locally.

Same configuration, opposite behaviour, and no error message in either case. If the code is installed locally the job will even succeed, and the results will be recorded against a machine that never ran them.

A host alias like hpc-tunnel is never on that list, so it is always treated as remote. Use one. This is tracked as #144.


If neither is allowed

Some centres permit no scripted ssh at all. For those, ECCE has a dummy submission mode: it prepares the input deck and the submit script and stops, telling you where they are. You copy them over, run the job yourself, copy the output back, and import it.

siteconfig/Machines ships a dummy machine for this. See #94.


Troubleshooting

ECCE still asks for a password. The shared connection is not being found. Check ssh -O check myhpc, and confirm the name you registered in ECCE is the alias and not the real host name — ECCE runs ssh <name>, so only the alias picks up your ~/.ssh/config entry.

The job says it started and nothing runs. That is a different problem; see #141, fixed in 8.13.2.

Something else. Run the diagnostic collector and attach its output to an issue:

curl -O https://raw.githubusercontent.com/FriendsofECCE/ECCE/main/packaging/ecce-diagnose
chmod +x ecce-diagnose
./ecce-diagnose

It records which machines you have registered, how their names resolve, and whether this host can ssh to itself — which is usually enough to see what went wrong. See #143.


A note on what has been verified

The localhost behaviour described above was confirmed by reading RCommand::isRemote() in the source. The connection-sharing route follows from ECCE invoking plain ssh <machine> with no options that would defeat it, and was contributed by a user running ECCE against a 2FA-protected centre this way — but it has not been tested against such a centre by the maintainers, who do not have one. If you try it, please say how it went on #94.