Hire AI engineers who can get it running.

There is the AI you can talk about and the AI you can get to run on a machine. Candidates fine-tune a model and stand up a retrieval pipeline in an isolated cloud environment, and what comes back is whether it works, not how well they described it.

lab terminal

$ docker compose up -d

3 services started

$ systemctl status api

active (running)

$ curl -s localhost:8080/health

{"status":"ok"}

Live environment

Running
  • vmlab-vm-01Ready
  • svcapi-serverReady
  • lbgatewayProvisioning

2h 12m left of the 3h limit

Talking about models and shipping one are different jobs

Without Scalyz

You ask about embeddings and transformers, and grade how fluent the answer is.
A strong reading of a paper tells you nothing about a broken CUDA install.
Anyone can describe fine-tuning. Few have made it converge.

With Scalyz

They stand up a retrieval pipeline on a real machine, and you see whether it returns anything.
The operational work, the part that actually breaks, is the part being tested.
The candidate fine-tunes a model in the environment, and the result either runs or does not.

What is at stake

The job is not the demo, it is the day the demo breaks

The hard part of an AI role is rarely the idea. It is making it run, and keeping it running, on infrastructure that fights back.

Where it breaks

Between notebook and server

Plenty of candidates can make something work in a notebook. The role is what happens when it has to run somewhere else.

What an interview misses

The operational half

Model intuition is easy to talk about and hard to verify. Getting the thing to actually run is easy to verify and hard to fake.

What you are hiring for

It runs on Monday

You are not hiring a description of a pipeline. You are hiring the person who gets it up and keeps it up.

Live in an afternoon

Three steps from the job post to a shortlist you can defend, with no screening calls.

Backend hiring, July

Open
6 invited · 4 completed

Closes Friday, 18:00

Invite the whole list at once

You pick the test and the deadline, then send your whole applicant list at once, by email, CSV, or a shareable link.

Live environment

Running
  • vmlab-vm-01Ready
  • svcapi-serverReady
  • lbgatewayProvisioning

2h 12m left of the 3h limit

Candidates work when it suits them

Each candidate gets their own machine, a real one, with a real task already loaded.

Camille Durand

DevOps assessment

84
Infrastructure as code86
Networking78
Incident recovery91

No integrity flags

The shortlist ranks itself

Scoring runs automatically against the task you set, so the ranking builds as candidates finish.

lab terminal

$ docker compose up -d

3 services started

$ systemctl status api

active (running)

$ curl -s localhost:8080/health

{"status":"ok"}

Mission

3h limit

Ship it behind a load balancer

  • DoneSet up the web server
  • DoneDeploy the API service
  • Route traffic through the load balancer
  • Prove the health checks pass

2 of 4 complete

We test the half of the job that actually runs

The published missions here are operational: fine-tuning a model, and building a retrieval pipeline with LangChain. Both run in a real environment on the candidate's own machine, and both are scored on what the system does at the end. We do not grade research or model theory, we grade whether it works.

  • Every mission runs in an isolated cloud environment, one per candidate
  • Scored automatically on whether the system runs at the end, not on an answer
  • You can import your own scenario as a .zip and run it yourself before you send it

The questions we get

Straight answers on how the tests, the scoring and the credits actually work.

What is a Scalyz lab, exactly?

A real virtual machine with a real mission, not a quiz. The candidate works with the same tools they'd use on the job, and the session is recorded so you can review the evidence later.

Who scores it?

Automatically. The checks run against what the candidate delivered, and you get a report broken down skill by skill, so no one on your team grades a thing by hand.

Can they cheat with AI?

A chatbot can suggest commands, but it can't run a live system for the candidate. The mission has to be done on the machine itself, and any cheating signals come with context, so you can judge them yourself.

What do I pay?

In credits: one per candidate. A credit is reserved when you send an invitation and released if it expires unused. You pay per candidate tested, never per user.

Can I bring my own exercise?

Yes. Upload your files as a .zip, rename the mission to match your role, and run it yourself for one credit first. What you tested is exactly what candidates get.

What is it like for candidates?

On their own time. They start when it suits them, work with real tools on a real machine, and finish in one sitting. It runs in English and French.

Test the AI work that has to run

Book a demo and we will run one of these missions on your own open role.

One credit per candidate, reserved on send, released on expiry