text similarity distinct method using spark?

Multi tool use


text similarity distinct method using spark?
I want to get a text similarity distinct method on 200 million differnent sentences using spark .Suppose I have 4 sentence that is
["Hi I heard about Spark","Hi I heard about Spark World",
"Logistic regression models ","Logistic regression goodmodels "]
I hope get the reuslt is
["Hi I heard about Spark", "Logistic regression models]
Since the first sentence is similar to second sentence and the third sentence is similar to the 4th sentence arrorcding to Levenshtein distance:https://rosettacode.org/wiki/Levenshtein_distance
How to achieve it efficiently using spark? Because the data is 200 million, I am hesitate to do cartesian
By clicking "Post Your Answer", you acknowledge that you have read our updated terms of service, privacy policy and cookie policy, and that your continued use of the website is subject to these policies.