text similarity distinct method using spark?

The name of the pictureThe name of the pictureThe name of the pictureClash Royale CLAN TAG#URR8PPP


text similarity distinct method using spark?



I want to get a text similarity distinct method on 200 million differnent sentences using spark .Suppose I have 4 sentence that is



["Hi I heard about Spark","Hi I heard about Spark World",
"Logistic regression models ","Logistic regression goodmodels "]



I hope get the reuslt is
["Hi I heard about Spark", "Logistic regression models]



Since the first sentence is similar to second sentence and the third sentence is similar to the 4th sentence arrorcding to Levenshtein distance:https://rosettacode.org/wiki/Levenshtein_distance



How to achieve it efficiently using spark? Because the data is 200 million, I am hesitate to do cartesian









By clicking "Post Your Answer", you acknowledge that you have read our updated terms of service, privacy policy and cookie policy, and that your continued use of the website is subject to these policies.

Popular posts from this blog

Arduino Mega cannot recieve any sketches, stk500_recv() programmer is not responding

Visual Studio Code: How to configure includePath for better IntelliSense results

C++ virtual function: Base class function is called instead of derived