-
Notifications
You must be signed in to change notification settings - Fork 28k
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
[SPARK-3097][MLlib] Word2Vec performance improvement #1932
Conversation
Jenkins, test this please. |
@@ -34,7 +34,7 @@ import org.apache.spark.mllib.rdd.RDDFunctions._ | |||
import org.apache.spark.rdd._ | |||
import org.apache.spark.util.Utils | |||
import org.apache.spark.util.random.XORShiftRandom | |||
|
|||
import org.apache.spark.util.collection.PrimitiveKeyOpenHashMap | |||
/** |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
add an empty line after imports
QA tests have started for PR 1932. This patch merges cleanly. |
QA results for PR 1932: |
Jenkins, test this please. |
QA tests have started for PR 1932 at commit
|
QA tests have finished for PR 1932 at commit
|
Jenkins, test this please. |
QA tests have started for PR 1932 at commit
|
QA tests have finished for PR 1932 at commit
|
LGTM. Merged into master and branch-1.1. Thanks! |
mengxr Please review the code. Adding weights in reduceByKey soon. Only output model entry for words appeared in the partition before merging and use reduceByKey to combine model. In general, this implementation is 30s or so faster than implementation using big array. Author: Liquan Pei <liquanpei@gmail.com> Closes #1932 from Ishiihara/Word2Vec-improve2 and squashes the following commits: d5377a9 [Liquan Pei] use syn0Global and syn1Global to represent model cad2011 [Liquan Pei] bug fix for synModify array out of bound 083aa66 [Liquan Pei] update synGlobal in place and reduce synOut size 9075e1c [Liquan Pei] combine syn0Global and syn1Global to synGlobal aa2ab36 [Liquan Pei] use reduceByKey to combine models (cherry picked from commit 3c8fa50) Signed-off-by: Xiangrui Meng <meng@databricks.com>
------------------ 原始邮件 ------------------ 主题: Re: [spark] [SPARK-3097][MLlib] Word2Vec performance improvement(#1932) — |
mengxr Please review the code. Adding weights in reduceByKey soon. Only output model entry for words appeared in the partition before merging and use reduceByKey to combine model. In general, this implementation is 30s or so faster than implementation using big array. Author: Liquan Pei <liquanpei@gmail.com> Closes apache#1932 from Ishiihara/Word2Vec-improve2 and squashes the following commits: d5377a9 [Liquan Pei] use syn0Global and syn1Global to represent model cad2011 [Liquan Pei] bug fix for synModify array out of bound 083aa66 [Liquan Pei] update synGlobal in place and reduce synOut size 9075e1c [Liquan Pei] combine syn0Global and syn1Global to synGlobal aa2ab36 [Liquan Pei] use reduceByKey to combine models
@mengxr Please review the code. Adding weights in reduceByKey soon.
Only output model entry for words appeared in the partition before merging and use reduceByKey to combine model. In general, this implementation is 30s or so faster than implementation using big array.