Tech

Google asked staff to spend time teaching its Bard chatbot to write like a human. AI researchers explain how it works.

Google CEO Sundar Pichai speaks during the keynote address of the Google I/O conference in Mountain View, Calif, May 2019.
Google CEO Sundar Pichai. Jeff Chiu/AP Photo
Read in app

Last week, Google internally kicked off "dogfooding," where employees across the organization were asked to spend two to four hours helping test Bard, its new artificial-intelligence chatbot for search. 

The unveiling of Bard came shortly after Microsoft announced a revamped version of its Bing search engine that incorporates the ChatGPT bot. It allows users to have a back-and-forth dialogue on just about any subject. Google took a slight reputational hit after it was discovered that Bard answered a question incorrectly. Similarly, as more people have tested the new Bing, they've run into problems with that engine's bot, like its propensity to behave combatively

Bots like Bard and ChatGPT work by getting trained on text written by humans so they can mimic them. That explains why Bing may sound somewhat emotional and unpredictable — a bot trained to act human will do so, faults and all.

These bots initially do much of their learning by ingesting large sets of training data. In addition, Bard's product lead, Jack Krawczyk, told staff in a memo that the company's own work found that adding high-quality responses to user queries "dramatically" improved the AI model's quality.

AI experts told Insider how Googlers might write the high-quality responses for Bard to improve its model. These experts have completed extensive study in the fields of AI and large language models.

Bots can learn in different ways

Krawczyk told staff to ask Bard questions about areas in which they had domain expertise, such as a favorite hobby. Then they were asked to evaluate Bard's answers to ensure they were what one would expect and of a reasonable length and structure. If an answer was too emotionally charged, factually wrong, or otherwise didn't make sense, employees could rewrite the answer and submit it to help train Bard's model.

To refine Bard, Google could implement a combination of supervised and reinforcement learning, Vered Shwartz, an assistant professor of computer science at the University of British Columbia, said.

Supervised learning is the first step, where the chatbot is fed human-written queries and responses until it learns how to write like a human. The company could then layer on top a reinforcement-learning model that would be trained with answers written by Googlers to help it understand which values the company wanted Bard's answers to exhibit, whether it be in terms of structure, tone, or other qualities.

That model would look at answers Bard produced, rejecting the bad ones and validating the good ones until the chatbot understood how it should behave. Essentially, "good" answers from Googlers would fine-tune the model.

The reinforcement model could teach Bard to be informative without speaking with emotion or otherwise pretending to be human. The first model learns fundamental writing skills, while the second would steer responses in the desired direction.

With enough-good answers to analyze, the reinforcement model would be able to learn what's appropriate and what's not, Zhou Yu, a computer-science professor at Columbia University, said.

Factual accuracy 

Google has been cautious about its rollout of chatbots, likely because of the near-term influence it could have on search margins and concerns of accuracy. It told employees to reject responses to questions where Bard tried to provide a user with advice on sensitive topics like finance or health, as the risk of incorrect answers is high.

While training will improve the quality of generated answers, Shwartz said she didn't think it would completely solve the issue of factual accuracy. Bard and ChatGPT have a tendency to "hallucinate," a term the industry has adopted to say the bots make things up. They will pull content from web pages and summarize them incorrectly at times.

"Bots are trained to produce humanlike text, not to be truthful," Shwartz said.

The industry has been working to address factual accuracy, with OpenAI releasing an update in January to improve its factuality across a variety of topics. At a conference about chatbots and AI in San Francisco this month, Anthropic CEO Dario Amodei said he believed chatbots would stop making up facts as the models improved. 

Read next

Thomas Maxwell is a former Big Tech reporting fellow covering major companies including Alphabet and Microsoft. He also writes about emerging technology trends, like the rise of generative AI.Thomas previously spent two years at Input, a technology news publication owned by Bustle Digital Group. Before that, he completed internships at PBS NewsHour and Yale University Press. Got a tip? Thomas can be reached via email at tmaxwell@jkmperu.com, Signal at +1 540.955.7134, or Twitter at @tomaxwell. Use a personal device.Here are some examples of his work:Inside Google's latest moves to lock down its AI research before ChatGPT eats its lunch: 'It's time to compete'Leaked messages show Googlers are taking out their frustrations over layoffs on its new Bard AI chatbotInside Reddit's path to an IPO, where employees see 'thrash' from constant pivots and say more managers may leave amid a flatteningGoogle is downsizing its contract workforce that supports YouTube shortly after one contractor team's union victoryGoogle's exclusive $15 billion deal with Apple is the missing piece that Microsoft's Bing needs to win searchGoogle contractors say they don't have enough time to verify correct answers from the company's AI chatbot and end up guessingThe ChatGPT and generative-AI 'gold rush' has founders flocking to San Francisco's 'Cerebral Valley'The top 14 most influential Instagram executives in 2023Developers are turning to GitHub Copilot, a ChatGPT-like tool that helps them write code. One startup VP says it helped him save 10% of the time he'd spend coding.