Menu
← Selected Fieldwork

When the AI wouldn't listen

The ask: "How does an AI-powered rental search compare to traditional platforms over a real, active apartment search?" Short answer was that the AI-agent model consistently ranked last during a longitudinal study.

RoleLead Researcher
Duration~2.5 Weeks

All proprietary information and sensitive platform data has been sanitized or otherwise obscured.

Context

Who was the client and what was their ask?


The client was an apartment search platform that served renters across the United States. In early 2024, their product team wanted to understand how users experienced their platform over time compared to competitors — including their new AI-powered rental search agent.

I was brought in as the lead researcher through Askable, a research operations platform, to design and run the study end to end. I owned the screener, participant recruitment and vetting, moderation across all study touchpoints, analysis, and final deliverable.

The participants were five renters actively searching for apartments in major metros: Austin, New York City, Miami, and Las Vegas. They ranged from first-time solo renters to a mother of two relocating from out of state.

↗ Click to expand

The client wanted to understand how renters experience their platform over time — not just in a single session, but across a real, active apartment search.

The ask: How does the AI-powered, conversational interface fare when compared to home search platforms with traditional filters?

Methodology

To answer their ask, I designed and ran a diary study. Participants used four platforms simultaneously while actively searching for a real apartment in a city they were moving to. They weren't simulating a search. The stakes were genuine.

The four platforms: Apartments.com, Zillow, Apartment List, and the AI-powered rental search tool. The AI tool's defining proposition was that it removed traditional search filters entirely. Instead of letting users set their own parameters, it asked questions and generated results on their behalf. The user's job was to respond to the AI. The AI's job was to find the apartment.

The original brief called for check-ins with each participant three times per week for thirty minutes each. I raised concerns about participant fatigue, and so I and the client settled on a lighter-touch cadence: one mid-point survey, one check-in call, one final interview. Enough structure to track change over time.

Data was collected across three touchpoints:

  • On Day 3, a structured survey

  • On Day 4, a 30-minute call

  • Day 7, a 45-minute final interview.

Each touchpoint informed the next.

The cities we chose to examine were Austin, New York City, Miami, and Las Vegas. And the people we chose for the study were searching for an apartment in the city they were actively moving to — meaning the geography wasn't assigned.

We wanted to understand how the experience played out in their real life.

Timeline

↗ Click to expand

What happened

The AI-Powered Search Agent platform required users to complete a longer onboarding questionnaire before seeing results . This was a deliberate design choice meant to surface more relevant matches. But they had no longitudinal evidence that the tradeoff was worth it. Did the upfront friction actually pay off over time? They didn't know.

That was the question they hired me to answer.


Now, a this point, the client had invested significantly in building a conversational AI-agent, but by the time it reached me, it hadn't been tested for good market fit. It had been built, shipped, and assumed to be good. (This is pretty common, though. Teams fall in love with what they've built, and the value of AI in 2024 often went unquestioned).

A few things happened during the actual research process. These were:

  • Fraudulent participants! We had screened somewhere around ~1300, scheduled 25 for onboarding. For Onboardings, I would do identity verification calls, and I was actully surprised by how many weren't who they said they were. In the end, we only had 5 users who were eligible.

  • Because of this, along with other participant scheduling mayhem, the scope ballooned from 20 hours to 104.5 hours.

  • Users abandoned the AI tool because it gave them no way to iterate on their own search. The tool had removed the filter UI entirely. The only path to adjusting results was the chat prompt, and when users entered a refinement, the agent frequently returned results in a different direction than they intended.

    With no filters to fall back on, users had no way to course-correct.

    By the second interview, most had stopped using it.

    ↗ Click to expand


The Results

↗ Click to expand

What I learned

If I could go back to Day 1 of this project, I'd tell myself a few things.

I. Flag the client provided research question.

The research question, "How does the AI-powered, conversational interface fare when compared to home search platforms with traditional filters?"

The problem with that question was that it was comparing two completely different things. It was like a race, except with three horses and one bicycle. I should've pushed back on the client-provided question and suggest a new question, "Does the AI agent help or hinder people in an actual search, and why?"

And if I had more resources and freedom to design the study, I would've done the following:

  1. Target 6-8 users.

  2. Track One: Don't run a comparison study on the AI-powered agent just yet. Instead, make the three platforms with search filters rank against each other.

  3. Track Two: VALIDATE the actual UX of the AI-agent. Did it get more accurate? Could users get what they wanted? I would've suggested this come before anything else.

  4. Expand the timeline from 7 days to 3 weeks to better reflect the average length of the typical home search.

    ↗ Click to expand


II. Hammer out the scope before you start.

Not a general agreement, a specific one.

Hours, deliverables, touchpoints, what happens if any of those change. A kick-off agreement that covers timeline and recruitment logistics but leaves room for a client to add three check-ins per participant per week is not a complete scope document.

I learned the difference the hard way.

III. Vet your recruitment infrastructure like you vet your participants.


Recruiting is consistently the hardest part of research, and it gets harder when you're in an unfamiliar environment with an unfamiliar panel.

I assumed the platform would handle what it said it would. It didn't. Fraudulent participants made it through a panel that was supposed to pre-screen them, and I caught the problem myself through identity verification calls. That work was never in scope and never compensated.

Before I use any recruitment platform again, I'll run a small pilot, check references from other researchers, and understand exactly where their responsibility ends and mine begins.

The research was solid and the findings were real and deeply impactful. But because of the recruitment issues, I did 104 hours of work when I agreed to only 20. Needless to say, I only need to learn this lesson once.

Appendix

Team: Askable

← All fieldwork
Built with love @ 2026 Leeza Dennis
LINKEDIN