Skip to main content
Start main content

Understanding and Modelling Conversation

Cover Story

TonyMcEneryNewly appointed to the PolyU Faculty of Humanities, Prof. Tony McEnery, Chair Professor of Corpus Linguistics and AI for the Humanities, is a leading scholar in linguistics. His work explores how everyday conversation unfolds and adapts in real time, and how large-scale language data can deepen our understanding of communication—paving the way for more natural and human-like AI interactions.

 

How does language change and adapt as we use it in conversation with another person? Conversation is an incredibly difficult task for both those speaking and listening – we process information on the fly and plan responses at very high speed. Yet we know relatively little about how conversation is organized. This has real world consequences - we will all have had the experience of interacting with a chatbot and feeling something is not quite right. The way it interacts, what it both does and does not do, feels slightly wrong. It can be off putting for users and can certainly drive people away from using intelligent agents. But fixing the problem is difficult because we have surprisingly little real conversational data to work with. Asking people to carry around a recorder and to record everything they say and hear is challenging. So often researchers opt for some artificial setting to collect such data, relying on ‘phone conversations or artificial conversations in a lab setting. But that comes at a cost – the interactions are less natural and it is the naturalness that we want. What can be done about this?

Cover Photo
I have been working since the early 1990s to deal with this problem for British English, where we now have some tens of millions of words of conversational data. But in that time a surprisingly stubborn problem has remained – we have relatively little data for the dominant variety of English, American English. In September at PolyU a major step towards addressing that gap will be unveiled. I have been working with teams at Lancaster University in the UK and Northern Arizona University in the US to develop a large dataset of naturally occurring conversations from across the US, gathering data from every state, as well as people from different ages and social classes. With AI assistance, the conversations we have gathered have been painstakingly transcribed to give us, for the first time, a publicly available dataset which can show us how American English conversation works. With over 11 million words of data available, we can start to work on better dialogue systems so that in the future interactions with chatbots can feel more natural and jar less. Getting the right data is often the key to building better systems – and with our new dataset the prospect of better modelling conversation will become a real prospect.

Your browser is not the latest version. If you continue to browse our website, Some pages may not function properly.

You are recommended to upgrade to a newer version or switch to a different browser. A list of the web browsers that we support can be found here