ResearchRemoteFull time
AI Engineer
Build Depot-1, the intelligence behind Link: how it reads an operation, reasons about it and decides where attention should go next.
About Wehand
Wehand is an AI research and deployment company building intelligence for logistics. Link, our operational intelligence platform, brings together the people, vehicles, routes, money, documents and performance data that delivery and logistics operators already have, and helps them understand what is happening, why, and what to do next. Depot-1 is the intelligence behind it.
Our customers run real operations: vans leaving depots before sunrise, drivers to pay, carriers grading them every week. The software we build has to be right, fast and trusted by people who do not have time for anything that is not.
The role
You will build the intelligence at the centre of Wehand. Depot-1 is how Link understands an operation: it answers questions in Ask Link, investigates what is driving performance and recommends the tasks that matter most.
This is applied AI work with consequences. An answer that sounds right but is wrong can cost an operator money, so the job is as much about measurement, guardrails and honest limits as it is about capability.
What you will do
- Design how Depot-1 combines models, tools, retrieval and an operation's own data to answer questions and recommend actions
- Build evaluation sets and pipelines that measure accuracy, reasoning and failure modes on real operational tasks
- Develop the context and memory systems that decide what Link knows about an operation and when it uses it
- Build guardrails so that Link says when it does not know, cites where an answer came from and keeps people in control
- Work on cost, latency and usage limits so that the intelligence stays fast and affordable at scale
- Contribute to our research into evaluation, uncertainty and more autonomous execution
What we are looking for
- Experience shipping LLM-powered features to real users in production
- Strong software engineering, ideally in TypeScript or Python
- Hands-on experience with prompting, tool use, retrieval and structured outputs
- A rigorous approach to evaluation, and the habit of measuring before believing
- Honest judgement about what language models are and are not reliable for
- Clear written English
Nice to have
- Experience with agents, planning or multi-step reasoning systems
- A background in machine learning research, statistics or forecasting
- Familiarity with logistics or operational data
How we work
Remote, with a small team and real ownership. There are no layers between you and the decisions, and the work you ship reaches operators within days, not quarters.
We write things down, we talk to customers directly, and we would rather ship something small that works than something large that might.