Add distributed lookup table design #9075

jacquesqiao · 2018-03-14T11:56:00Z

No description provided.

wangkuiyi · 2018-03-14T23:39:23Z

Thanks to @jacquesqiao for this design. After talking with Xu laoshi, I put our comments in jacquesqiao#4. Please review it. Thanks!

Update distributed lookup table design doc

typhoonzero · 2018-03-15T02:08:00Z

Related: #9068

typhoonzero

Basically LGTM, just some questions.

typhoonzero · 2018-03-15T02:46:40Z

doc/design/distributed_lookup_table_design.md

+
+<!--
+Note: please update the following URL when update this digraph.
+<img src='https://g.gravizo.com/svg?


One awesome tool for pasting figures!

typhoonzero · 2018-03-15T02:51:41Z

doc/design/distributed_lookup_table_design.md

+
+- Pro: parameter servers do not retrieve W(x) from the storage
+  service, thus saves half network communication.
+- Con: the storage service needs to be able to run the optimization


Does this means we actually have two kinds of servers when doing training:

original parameter server

storage server can run optimization

yes, storage server can run optimization will only handle large-scale embedding table, other parameters still use the fluid optimization operators.

Yes, if we go this way. But let us go the other way first.

jacquesqiao · 2018-03-15T03:01:13Z

doc/design/distributed_lookup_table_design.md

+### Storage Service Doesn't Optimize
+
+In this design, we use highly-optimized distributed storage, e.g.,
+memcached, as the storage service, and we run the optimization


@wangkuiyi After discussing with @ldmod LiDong, We think that a standalone distributed KV service like Memcached maybe not a good solution. It's better that parameter and the corresponding optimization happened at the same place. For large scale model training and optimization, we need to use the asynchronous update, in this condition, we need to make sure every optimizer should update the latest parameter, even when the gradient is calculated by the old parameter. If we use a standalone KV service, we need to add a lock to the parameter when some optimizer is updating it. If the parameter and optimization op is at the same place, the solution will be very simple, for example, use one thread to do the optimization.

But still, we can have two option, one is using the current optimization operators, but the parameter will be maintained by the pserver. The other one is using a distributed Storage Service which can do Optimization.

Discussed with @Yancey1989 yesterday that we agree with this idea, mixing two servers (original pserver and embedding table pserver) is a perfect way to embed an external parameter server:

Add design doc for lookup remote table in Fluid #9068 is trying to describe how to implement this opensource.

embed Abucus parameter server for advanced industry model training

we can carry on this two methods at the same time, then both opensource users and industry users can have their choice.

To using distributed Storage Service, PServer would also communicate with the storage, this would make the double traffic than using embedding table pserver. And I agree with @typhoonzero , we can carry on this two methods at the same time.

Let us have a baseline solution as early as possible. @jacquesqiao @ldmod.

@wangkuiyi I agree. We can implement a baseline solution without Memcached.

reyoung · 2018-03-15T04:41:28Z

How to implement sparse regularization?

helinwang · 2018-03-15T17:50:12Z

doc/design/distributed_lookup_table_design.md

+
+## Conclusion
+
+Let us do the "storage service does not optimize" solution first, as a


Sorry I don't understand why we need to use memcached given that our current parameter server implementation should be faster in key lookup.

memcached is a general purpose in-memory key value store, from its website description:

Memcached is an in-memory key-value store for small chunks of arbitrary data (strings, objects) from results of database calls, API calls, or page rendering.

We only need a special case in-memory key value store (few keys, large values). Currently our parameter server (recv operator in fluid) does exactly this. In terms of the key lookup speed, our parameter server should be faster or at least same speed comparing to memcached.

The design doc mentioned:

Let us do the "storage service does not optimize" solution first

Our parameter server already does this (and does optimization as well), why we have to make another implementation with memcached again, is it because of performance reasons?

@helinwang I am not so clear about the table lookup implementation currently, for a large scale lookup table, the keys may be discrete in a very big range, so we need a key-value lookup module.

@jacquesqiao

Will discuss offline with @jacquesqiao

jacquesqiao · 2018-03-19T06:37:33Z

Will issue a detailed design in another PR.

jacquesqiao added 3 commits March 14, 2018 19:43

update

8c67fff

update format

0baf4e1

typo

267ffc2

jacquesqiao changed the title ~~Add distributed lookup table design~~ [wip]Add distributed lookup table design Mar 14, 2018

Update distributed lookup table design doc

4c33a10

Merge pull request #4 from wangkuiyi/yi-lookup-op

bc611f0

Update distributed lookup table design doc

jacquesqiao changed the title ~~[wip]Add distributed lookup table design~~ Add distributed lookup table design Mar 15, 2018

jacquesqiao requested a review from wangkuiyi March 15, 2018 01:34

update digraph format for markdown

5cf5ad3

update graph url

c95156e

jacquesqiao mentioned this pull request Mar 15, 2018

Add design doc for lookup remote table in Fluid #9068

Merged

Yancey1989 added this to To do in distributed lookup table Mar 15, 2018

jacquesqiao requested review from Yancey1989 and typhoonzero March 15, 2018 02:32

typhoonzero approved these changes Mar 15, 2018

View reviewed changes

jacquesqiao commented Mar 15, 2018

View reviewed changes

jacquesqiao requested review from helinwang and typhoonzero March 15, 2018 03:05

helinwang previously requested changes Mar 15, 2018

View reviewed changes

jacquesqiao merged commit d7d0c1e into PaddlePaddle:develop Mar 19, 2018

distributed lookup table automation moved this from To do to Done Mar 19, 2018

jacquesqiao mentioned this pull request Mar 19, 2018

Support Distribute Lookup Table #9211

Closed

15 tasks

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Add distributed lookup table design #9075

Add distributed lookup table design #9075

jacquesqiao commented Mar 14, 2018

wangkuiyi commented Mar 14, 2018

typhoonzero commented Mar 15, 2018

typhoonzero left a comment

typhoonzero Mar 15, 2018

typhoonzero Mar 15, 2018

jacquesqiao Mar 15, 2018

wangkuiyi Mar 15, 2018

jacquesqiao Mar 15, 2018 •

edited

Loading

jacquesqiao Mar 15, 2018

typhoonzero Mar 15, 2018

Yancey1989 Mar 15, 2018

wangkuiyi Mar 15, 2018

jacquesqiao Mar 15, 2018

reyoung commented Mar 15, 2018

helinwang Mar 15, 2018 •

edited

Loading

jacquesqiao Mar 16, 2018

jacquesqiao commented Mar 19, 2018


		## Conclusion

		Let us do the "storage service does not optimize" solution first, as a

Add distributed lookup table design #9075

Add distributed lookup table design #9075

Conversation

jacquesqiao commented Mar 14, 2018

wangkuiyi commented Mar 14, 2018

typhoonzero commented Mar 15, 2018

typhoonzero left a comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

jacquesqiao Mar 15, 2018 • edited Loading

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

reyoung commented Mar 15, 2018

helinwang Mar 15, 2018 • edited Loading

Choose a reason for hiding this comment

Choose a reason for hiding this comment

jacquesqiao commented Mar 19, 2018

jacquesqiao Mar 15, 2018 •

edited

Loading

helinwang Mar 15, 2018 •

edited

Loading