We use cookies — so the site works and to count visits. By pressing "Got it" you also allow Webvisor session recording. Details and how to opt out
Dwelly
Data Analyst (AI agents)
11 days ago
Contacts
Reach out directly about this role
Details
Grade
Senior
Work Format
Remote
Employment
Full-time
English Level
B2 - Upper-Intermediate
Average salary for this role
By job title
110,000USDper year
85,000145,000
A cover letter in a minute
AI writes it from your resume. You only have to send it.
Description
We are digitizing and optimizing apartment rentals in the UK. Through a great product and optimal business processes, we are increasing the profitability of real estate agencies from an average of 10% to 40%.
This role involves researching, evaluating, and improving our core product: autonomous AI agents, through which all property management processes flow, from collecting payments from tenants to solving their plumbing issues. The agents themselves communicate with tenants, landlords, and vendors, make decisions within their authority, and only involve a human when absolutely necessary.
An example of a process managed by an agent
One month before the electrical safety certificate expires, the agent initiates the inspection and renewal process. They send inspection requests for a specific address to several vendors and collect quotes from those willing to undertake the work. They select the best vendor and provide them with the tenant's contact details to arrange a visit date. They monitor the status: receive confirmation and the visit date from the vendor. On the visit date, they verify that the vendor has submitted the inspection results. If not, they send a reminder. If the results indicate repairs are needed, they obtain repair cost quotes from several vendors. They agree on the best vendor and cost with the landlord. They provide the tenant's contact to the vendor. At the end, they issue and pay all invoices to all parties and update the certificate details.
Previously, humans handled all of this. The agent forgets nothing and acts instantly upon receiving the necessary information. This not only increases the efficiency of human labor but also significantly improves the quality of service for the tenant. However, autonomy has a downside: an agent can make mistakes confidently and silently. Therefore, our main analytical question is: how do we know if an agent is working well, and how can we make it work better?
Why Analytics is Needed Here
How can we determine if a new agent version (prompt, model, tools) is better than the old one before it goes into production? What offline acceptance metrics should we use? How many cases are needed in the evaluation set to ensure the difference between versions isn't just noise? How do we collect an evaluation set that reflects reality, not just the happy path? Where do we find rare but costly cases: a leak on a Friday evening, a tenant who has stopped responding, a landlord who disagrees with the estimate?
Can we run the agent through a simulation where the tenant, landlord, and vendor are also LLMs? To what extent do the simulation scenarios cover the real distribution of cases? And does the simulation result predict what will happen in production?
Can we trust an LLM-as-a-judge? How much do its assessments align with human annotation, where does it systematically err, and how can we measure this?
When an agent calls for human assistance, how justified is it? How many cases were there when help was needed, but the agent didn't call? How much time does a human take to react? Can the agent be improved so it can handle cases on its own?
What are the most frequent and most costly agent errors? What should be fixed first: the prompt, the tools, the data, or the process itself?
How do we understand the funnels of all communications (from all parties) and the timings of all processes to improve service quality? What constitutes vendor quality? Time spent booking a visit? Time until the visit date? What if a vendor responds to an email in minutes but schedules the repair only after 6 days? What if the delay is due to the tenant's initiative (they couldn't make it earlier) and not the vendor's? What signal should we look at to understand this?
How should we act if a client calls by phone and we need to intervene in an autonomous flow? How many employees are now needed to handle the previous workload? And how does the increased autonomy of agents translate into margin?
Should weekends be included in task timing calculations? And non-working hours in general? Do agents truly cover all property management scenarios? And what do coordinators do outside the system, communicating by phone?
Why It's Cool
Great founders with whom there's much to learn. A large, but compact and well-capitalized market. One of the few tasks where autonomous agents perform real operational work in production, and the quality of their work is directly reflected in profits. International experience; you'll hone your British English. Work with me (though there are differing opinions on that =)
We hire only seniors (your colleagues will be the same).
Requirements
Seeking a versatile master with a "Senior Analyst" level (requirements for levels).
Solid applied statistics and experimental design: confidence intervals, power, comparing versions on small samples.
Understanding of how LLM systems break and a desire to deeply understand agent evaluation.
Experience with evaluations, LLM-as-a-judge, or simulations will be a significant plus.
Remote work not from the Russian Federation/Republic of Belarus.
English sufficient for work communication.
Conditions
We pay market rates in pounds/dollars/euros (you must be able to receive payments in foreign currency).
Remote, time zone +/- London.
For successful candidates, we can arrange a work visa for the UK or a country in the EU through Deel.