When Experimentation Leads to Innovation: Machine Learning at FullContact | FullContact

When Experimentation Leads to Innovation: Machine Learning at FullContact

Paris Mitton February 7, 2018

At FullContact, we’re always experimenting with new technologies and techniques. Machine learning has come into vogue as of late, and has shown some impressive results within our company and without. Recently, we had an opportunity to apply some machine learning to improve our handling of job title data we find throughout the web. The choice, however, was not an easy one. To someone unfamiliar with the technology, machine learning can seem complex, foreign, expensive, and hard. Training data, precision and recall, neural networks, scary-sounding academic papers — it’s overwhelming. Given such a large investment, is it really the right choice? In this post, I’ll try to explain our thought process at FullContact on solving a tough problem, and how we decided to use machine learning.

The Problem

FullContact handles job title data in a variety of forms. Understanding job titles at a deep level is critical for contact search, analytics, and formatting for use in our APIs and consumer products. One problem that we had historically not solved was department classification. Given a job title, we needed to classify it into a corporate department.

For instance, a job title input of “software engineer” should return a department classification of “engineering.” “Controller” should return the “finance” department. “Customer Success Manager” is “Customer Support.

This solution may seem simple, but it quickly becomes complicated after further thought. For instance, what rules could we use to classify a job title into the “engineering” department? Maybe if the job title contains the word “engineer”? This works for “software engineer” but fails for “programmer,” which should also be an engineering job. This approach classifies “support engineer” as “engineering,” when at many companies that position is a customer support job. The story is similar for many other professions, and simple rules like “contains the word engineer” quickly become a quagmire of logic and exceptions.

This was the state our department classification methodology. It worked for common job titles, but it was painful to maintain and had undefined behavior for job titles outside the norm. Addressing bugs in this system was usually “add another special case,” because fixing the underlying logic was either not useful or too complex.

The Challenge

After talking about the problem, we decided that it was worth trying a prototype using machine learning. This was not a decision made lightly. We were well aware of the drawbacks of machine learning:

These factors made machine learning a difficult task to take on. Yet, we also understood that machine learning could potentially solve our problem more elegantly than any “bag of rules” approach. There’s a certain class of problems where machine learning does very well:

After reviewing this list, we decided that our problem fit the machine learning use case, and we went to work on a prototype. Going from nothing to a machine learning pipeline taught us many lessons through a process of trial and error. For instance:

Following these general guidelines, our application of machine learning definitely paid off. We ended up with an algorithm that correctly classified the department of a job title in 88 percent of our test set (and likely much more than that in production contexts). The algorithm correctly classified job titles that we had never considered.

How it Works

To understand the power of our machine learning classifier, we need explain how the job title department process works. We use a program called word2vec to consume a large quantity of job title and job description data, observing how and where words appear near each other. The data is then used to train a neural network to assign a vector to every word in the dataset. The vectors have a variety of special characteristics, but the most obvious one is that words that are similar to each other have similar vectors (“running” and “sprint” are more similar to each other than “sprint” and “cake”).

With our newly created data set, we have a metric for similarity between words. When we want to classify a job title, we combine the vectors tied to each word in a given job title to create an associated vector. We then compare the job title vector with programmer-defined vectors that represent each department (the vector for the department “engineering” might be close to the vectors for the words “engineer,” “programmer,” and “development”). The department vector that is the closest to the job title vector is considered the correct classification.

The Results

The algorithm that results from this machine learning approach lead to behavior that is much more elegant and nuanced than previous methodologies. There are some straightforward examples:

None of the answers came from having explicit rules. The associations were learned automatically by looking at how often words occur near each other from our training data. This saves programmers huge amounts of work considering every possible title, domain, and department, as well as the many ways people commonly express these job titles. It is resilient to minor changes in expression:

The resilience is due to the fact that low-signal words such as “interim” are weighted less in determining the results. What is really impressive is the ability of the model to learn associations for rare words:

That said, the results are not perfect.

Many of these examples are understandable — “product outreach coordinator” contains “product,” which in this case actually isn’t the product department. Others, like “remodeling consultant,” are a bit stranger, and unfortunately one of the drawbacks of our approach is that debugging bad word associations is difficult.

Overall, machine learning can be an incredibly powerful tool, but it has to be used in the right context. Its cost, complexity, and propensity for generating tech debt makes it no small undertaking. However, some problems, especially ones with complex and ambiguous solutions, can be solved more elegantly by machine learning than any other approach. Armed with the relevant expertise, a test set, and the willingness to grind for incremental improvements, the results can be quite impressive.