The Wait After Submit. Why Dodona is faster now. #179
bmesuere
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
You hand in a solution and the page says it is running. Three separate things have to happen before you see a result, and over the summer we made all three faster. A typical exercise now answers in about two seconds instead of six, and the slowest kind answers in three seconds instead of fifteen.
What are you waiting for?
Pressing hand in starts a chain. Your submission is put on a queue and waits for a worker machine to pick it up. That worker builds a sandbox, copies the judge into it, and runs your code inside a container. When it finishes, the result is in the database, but your browser does not know yet: the exercise page checks the server every so often, and only shows the result at the next check.
Note
Most numbers below are for two judges. A judge is the program that decides whether a solution is correct, and each exercise names the one it uses.
TESTed is the one we recommend today. It is language-independent: the same test suite works for Bash, C, C++, C#, Haskell, Java, JavaScript, Kotlin, Python and TypeScript. It was also by far the slowest.
The Python judge is the older single-language one. It is no longer recommended for new exercises, but it still carries the most submissions on Dodona by a wide margin. Both judges run Python exercises, so the numbers are split by judge rather than by language.
The other judges are covered further down.
Wait 1: The queue
A worker asks the database for new work on a fixed interval. When workers are idle, that interval is the whole wait: with a handful of workers asking every few seconds, your submission sits there until one of them next looks.
At the end of July we swapped the queue library, Delayed Job, for Solid Queue. Solid Queue is the more modern of the two and polls the database far more cheaply, so we could afford to look more often: first every second, then every half second. The median wait for a worker went from 553 to 102 milliseconds.
Wait 2: The judge
Judging a submission takes three steps: preparing a sandbox, running your code inside it, and cleaning up afterwards. Preparing means copying the exercise and the judge into a fresh directory and creating the Docker container that will run them.
Each judge lives in its own git repository, and Dodona keeps a copy on shared network storage. Every file copied from there costs a network round trip, which made the time to prepare a submission follow the number of files in the judge almost exactly.
TESTed is the extreme case, because it carries an implementation for every language it supports. Most of those files are not the implementations, though: 461 of its 596 files are the judge's own test suite, which is not needed while judging, and every TESTed submission spent about five seconds copying them. The other judges had the same problem in smaller amounts: the Java judge copied 224 redundant files and the HTML judge 69.
We made three changes to shorten this preparation.
A judge can now list what Dodona should leave out when it syncs the repository, in the same spirit as
.dockerignore. Every judge on Dodona has such a list now.The files that remain are no longer fetched one by one. Each judge is now stored as a single archive, built once when the judge changes and unpacked per submission, so the worker does one network read instead of hundreds.
Finally, the Docker container is created while those files are being copied rather than afterwards, and it is removed in the background once your result is in.
The preparation phase of a TESTed submission went from an average of 7.0 seconds to 0.6, and of a Python-judge submission from 0.8 to 0.5. The total time on the worker followed: the weekly average for TESTed fell from 12 seconds to 3.1, and for the Python judge from 3.9 to 2.2. That total is what is shown on the chart below.
These changes are not specific to the Python judge or TESTed: they apply to every judge on Dodona. Counting everything the worker does apart from running your code, they are worth roughly half a second to a second per submission, whatever language it is in.
Wait 3: Showing the results
Once the judge is done the answer sits in the database, but the exercise page does not know that yet. It therefore keeps polling the server. Until September it polled after 1, 3, 6, 10 and 15 seconds and then every five, so a result could only ever appear at one of those moments: a Python submission that was ready after 3.1 seconds was shown to you at 6.
The page now polls every half second for the first ten seconds, then every second, then every five. Polling that often only became affordable with the new hand-in flow. 78% of results are now on screen within three seconds and 90% within four; before, 69% took at least six seconds.
Adding it up: four seconds, 2.7 million times
All three waits are shorter now. A worker notices your submission within half a second, the sandbox is ready in a fraction of the time it used to take, and the page shows the result almost as soon as it exists. On a Python-judge exercise you get your result about four seconds sooner than in July; on a TESTed exercise, about twelve.
What is left is mostly the judge running your code, and that is where we are looking now. We are tweaking the memory usage of the judges and are migrating the worker servers from machines with spinning disks to ones with local SSDs. We're also looking at swapping out pylint for Ruff which is much faster. Each of those will shorten the judging step itself rather than the waits around it.
All reactions