Case study · AI & Knowledge · Property Technology
AI Property Assistant
Conversational property search, grounded in real listing data.
01The question
Could someone ask a property website “three bedrooms, near a station, with a garden?” - and get an answer drawn from actual listings, not invented by a model?
The problem: property information was spread across listing pages in inconsistent formats, so buyers have to click through them one by one to find a match. This proof of concept explored whether an assistant could answer their questions directly from real listing data - combining a data pipeline, AI-assisted cleansing, hybrid search and a conversational interface.
02How it works
Six stages, from raw web page to grounded answer.
01 Collect
Scheduled ingestion
A timer-triggered function gathers public listing pages overnight, following pagination and respecting robots.txt, with request pacing, retries with backoff and a record of anything that fails.
- Azure Functions
- Python
- Blob Storage
02 Clean
AI-assisted cleansing
Raw listings are inconsistent. A language model extracts each one into a strict JSON schema. Where information is missing it's left empty rather than filled with plausible-sounding text.
- Azure OpenAI
- JSON schema
- SQL logging
03 Index
Hybrid search index
Clean records are indexed for hybrid retrieval: keyword matching for exact details like postcodes and prices, vector search for meaning like “quiet” or “good for commuting”.
- Azure AI Search
- Vector search
04 Answer
Grounded responses
For each question, the API retrieves the most relevant listings and passes them to the model with a property-first prompt that forbids inventing facts. Answers stream back in real time.
- Azure OpenAI
- Next.js API
- Streaming
05 Experience
Embeddable assistant
A conversational interface with live property search, built to drop onto an existing website as a widget - an overlay on desktop, full screen on mobile.
- Next.js
- React
- TypeScript
06 Platform
Infrastructure as code
The environment is defined in Bicep - App Service, Functions, Storage, AI Search, Azure OpenAI and Key Vault - so it can be reviewed, repeated and rebuilt.
- Bicep
- Key Vault
- App Service
03Engineering decisions
The details that make it trustworthy.
- 01
Fix the data before blaming the model
Early ingestion runs surfaced empty titles, placeholder addresses, energy ratings stored as URLs and duplicated highlights. Rather than hoping the model would paper over it, these were fixed at source and validated - an energy rating is only accepted if it's actually A to G.
- 02
Two kinds of question
Someone viewing a listing might ask about that home, or about something else entirely. The assistant distinguishes listing questions (“does this one have parking?”) from discovery questions (“anything near the station?”) so it doesn't anchor on the wrong property.
- 03
Empty beats invented
The cleansing and answering steps are both instructed to prefer “not known” over a confident guess. For a property search, a missing detail is an inconvenience; a made-up one is a problem.
- 04
Designed to embed, not just to demo
Rather than a standalone demo, the assistant was designed as an embeddable widget with a lightweight loader script, so it could sit on an existing site without a rebuild.
04What it demonstrates
Capability across the whole stack
- Retrieval-augmented generation over structured and unstructured data
- Resilient data ingestion and AI-assisted data cleansing
- Hybrid keyword and vector search
- A full-stack, embeddable conversational application
- Azure infrastructure defined as code, with secrets in Key Vault
Status, honestly
This was only ever a proof of concept, built and tested against real, publicly listed property data. It hasn't been deployed to production or tested with real users, so there are no outcome metrics - and we won't invent them. It's here as evidence of what we can build end to end, not a sign we only work in property.
05Where else this applies
Built for property. Not only for property.
The same questions, again
Staff answering the same customer questions every day - about stock, services, bookings or policies - when the answers already exist.
Information nobody can find
Policies, manuals, spreadsheets and records spread across folders and systems. The real search engine is the one person who knows.
Copying between systems
Someone re-typing information from emails, forms and documents into another system - every day, by hand.
Reports that take hours
Numbers pulled from several places and rebuilt in a spreadsheet every week, instead of assembled automatically.
Have information your customers or team struggle to search?
Start a conversation