Dwelly — Senior Analyst (AI Agents)
We are digitizing and optimizing apartment rentals in the UK. Through a great product and optimal business processes, we are increasing the profitability of real estate agencies from an average of 10% to 40%.
This role involves researching, evaluating, and improving our core product: autonomous AI agents, through which all property management processes are handled, from collecting payments from tenants to resolving their plumbing issues. The agents communicate with tenants, landlords, and suppliers themselves, make decisions within their authority, and only involve a human when absolutely necessary.
An example of a process handled by an agent
- One month before the electrical safety certificate expires, the agent independently initiates the inspection and renewal process.
- Sends inspection requests for a specific address to several suppliers and collects quotes from those willing to take the job.
- Selects the best supplier and provides them with the tenant's contact information to arrange a visit date.
- Monitors the status: receives confirmation and visit date from the supplier.
- On the visit date, verifies that the supplier has sent the inspection results. If not, sends a reminder.
- If the results indicate defects requiring repair, obtains repair cost quotes from several suppliers.
- Agrees on the best supplier and cost with the landlord.
- Provides the tenant's contact information to the supplier.
- Finally, issues and pays all invoices to all parties and updates the certificate information.
Previously, humans handled all of this. The agent forgets nothing and acts instantly upon receiving the necessary information. This not only increases the efficiency of human labor but also significantly improves service quality for the tenant. However, autonomy has a downside: an agent can make a confident and silent mistake. Therefore, the main question for our analytics is: how do we know if the agent is working well, and how do we make it work better.
Why Analytics is Needed Here
- How can we determine if a new version of the agent (prompt, model, tools) is better than the old one before it goes into production? What offline acceptance metrics should be used? How many cases are needed in the evaluation set to ensure the difference between versions is not just noise?
- How to build an evaluation set that reflects reality, not just the happy path? Where to find rare but costly cases: a leak on a Friday evening, a tenant who has stopped responding, a landlord who disagrees with the estimate?
- Is it possible to run the agent through a simulation where the tenant, landlord, and supplier are also LLMs? To what extent do simulation scenarios cover the real distribution of cases? And does the simulation result predict what will happen in production?
- Can we trust LLM-as-a-judge? How closely do its assessments match human labeling, in what areas does it systematically err, and how can this be measured?
- When the agent calls for human assistance, how justified is it? How many cases were there where assistance was needed, but the agent didn't ask? How much time does a human reaction take? Can the agent be improved to handle these situations independently?
- Which agent errors are most frequent and which are most costly? What should be fixed first: the prompt, the tools, the data, or the process itself?
- How to understand the funnels of all communications (from all parties) and the timings of all processes to improve service quality?
- What constitutes supplier quality? Time spent booking a visit? Time until the visit date? What if the supplier responds to emails within minutes but only schedules repairs for 6 days later? What if the delay is due to the tenant's initiative (they couldn't do it earlier), not the supplier? What signal should we look at to understand this?
- How to act if a client calls by phone and intervention is needed in the autonomous flow?
- How many employees are now needed to handle the previous workload? And how does the increase in agent autonomy translate into margin?
- Should weekends be included in the calculation of task timings? And non-working hours in general?
- Do the agents truly cover all property management scenarios? And what do coordinators do outside the system, communicating by phone?
Why It's Cool
- Great founders with whom there is much to learn.
- A large, yet compact and well-capitalized market.
- One of the few tasks where autonomous agents perform real operational work in production, and the quality of their work is directly reflected in financial results.
- International experience, you'll improve your British English.
- We only hire seniors (your colleagues will be of the same caliber).
Requirements
- Seeking a versatile master with the level of "Senior Analyst".
- Solid applied statistics and experiment design: confidence intervals, power analysis, comparing versions on small samples.
- Understanding of how LLM systems break and a desire to deeply dive into agent evaluation. Experience with evaluations, LLM-as-a-judge, or simulations will be a big plus.
- Remote work not from RF/RB.
- English that allows for professional communication.
Conditions
- Competitive pay in pounds/dollars/euros (you must be able to receive payments in foreign currency).
- Remote work, time zone +/- London.
- For successful candidates, we can arrange a work visa in the UK or an EU country via Deel.