You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
We introduce Florence-2, a novel vision foundation model with a unified,prompt-based representation for a variety of computer vision andvision-language tasks. While existing large vision models excel in transferlearning, they struggle to perform a diversity of tasks with simpleinstructions, a capability that implies handling the complexity of variousspatial hierarchy and semantic granularity. Florence-2 was designed to taketext-prompt as task instructions and generate desirable results in text forms,whether it be captioning, object detection, grounding or segmentation. Thismulti-task learning setup demands large-scale, high-quality annotated data. Tothis end, we co-developed FLD-5B that consists of 5.4 billion comprehensivevisual annotations on 126 million images, using an iterative strategy ofautomated image annotation and model refinement. We adopted asequence-to-sequence structure to train Florence-2 to perform versatile andcomprehensive vision tasks. Extensive evaluations on numerous tasksdemonstrated Florence-2 to be a strong vision foundation model contender withunprecedented zero-shot and fine-tuning capabilities.
URL
Affiliations
Abstract
Translation (by gpt-3.5-turbo)
既存の大規模なビジョンモデルは転移学習に優れていますが、単純な指示でさまざまなタスクを実行することに苦労しています。これは、さまざまな空間の階層と意味の粒度の複雑さを扱う能力を必要とするからです。
Florence-2は、テキストプロンプトをタスクの指示として受け取り、キャプショニング、オブジェクト検出、グラウンディング、セグメンテーションなどの望ましい結果をテキスト形式で生成するように設計されています。
このマルチタスク学習のセットアップでは、大規模で高品質な注釈付きデータが必要です。そのため、私たちはFLD-5Bを共同開発しました。これは、自動化された画像注釈とモデルの改善の反復戦略を用いて、1億2600万枚の画像に対して54億の包括的な視覚注釈を行ったものです。
私たちは、Florence-2を多目的かつ包括的なビジョンタスクを実行するためにシーケンスツーシーケンス構造を採用しました。数多くのタスクでの徹底的な評価により、Florence-2が前例のないゼロショットおよびファインチューニングの能力を持つ強力なビジョン基盤モデルの候補であることが示されました。
Summary (by gpt-3.5-turbo)