In this article I would like to present several metrics to calculate the similarity between sets of items. I’ve been analyzing diverse metrics as part of an investigation I’m doing to improve the targeting of digital advertising campaigns. These set similarity metrics are very useful to address the problem of audience expansion.
The basic idea is: we start with a set of users that have engaged with the ad, for example clickers. Then we try to find other similar sets of users that we can target. We have the expectation that, because of their similarity, the targeted users will also engage positively with the ad. These similar users become our expanded audience.
Jaccard Similarity
The Jaccard Similarity is defined as the size of the intersection divided by the size of the union of the sets.
Given two sets, A and B, the Jaccard Similarity is defined as:
The Jaccard Similarity ranges between zero and one.
Also called: Jaccard index, Intersection over Union
For more information: Jaccard Similarity
Sorensen Coefficient
The Sorensen Coefficient equals twice the number of elements common to both sets divided by the sum of the number of elements in each set.
Given two sets, X and Y, the Sorensen Coefficient is defined as:
The Sorensen Coefficient ranges between zero and one.
Also called: Sorensen–Dice index, Sorensen index, Dice’s coefficient
For more information: Sorensen Coefficient
Tversky Index
For sets X and Y, the Tversky Index is given by:
Note that are parameters of the Tversky Index.
The Tversky Index ranges between zero and one.
The Tversky Index can be seen as a generalization of the Jaccard Similarity and the Sorensen Coefficient:
- Setting produces the Jaccard Similarity.
- Setting produces the Sorensen Coefficient.
Tversky measures with are of special interest.
For more information: Tversky Index
Overlap Coefficient
The Overlap Coefficient is defined as the size of the intersection divided by the size of the smaller of the two sets.
For sets X and Y, the Overlap Coefficient is given by:
If set X is a subset of Y or the converse then the Overlap Coefficient is equal to 1.
Also called: Szymkiewicz–Simpson Coefficient
For more information: Overlap Coefficient

I’ve observed in recent years that many professionals became obsessed with networking. These are the people who constantly participate in conferences, meetups and other kinds of networking events. They like to collect business cards and feel excited whenever they connect to other professionals on LinkedIn.
In general the goal of networking is to create new opportunities. These may be business opportunities, partnership opportunities or job opportunities. But in my opinion what really creates new opportunities is our reputation.
Lots of people today dream about having their own startup company. These aspiring entrepreneurs are constantly looking for ideas, and in general they believe that a good startup idea must be very innovative, and better yet if they can be the
The work of a Software Engineer is to analyze a problem, think about a good solution, design it, implement it and then test it. Software Engineers do problem-solving. The work of the Software Engineer ends when he finishes implementing the solution for the problem.
Software Engineering teams are mostly busy implementing new features. Data Science teams are mostly busy running experiments.
Last week 
In decision-making, we try as much as we can to impact future results based on our previous experience. But every decision is also based on the assumption that its positive consequences will be greater than any possible negative outcome. Therefore, we should treat decisions as experiments. We should always preserve the option of reconsidering our decisions in the future, when we will be able to measure their real impact.