forked from docs/doc-exports
Reviewed-by: Pruthi, Vineet <vineet.pruthi@t-systems.com> Co-authored-by: Su, Xiaomeng <suxiaomeng1@huawei.com> Co-committed-by: Su, Xiaomeng <suxiaomeng1@huawei.com>
130 lines
15 KiB
HTML
130 lines
15 KiB
HTML
<a name="dli_08_15050"></a><a name="dli_08_15050"></a>
|
|
|
|
<h1 class="topictitle1">Result Table</h1>
|
|
<div id="body0000001712361274"><div class="section" id="dli_08_15050__section77915506517"><h4 class="sectiontitle">Function</h4><p id="dli_08_15050__p1529893619479">This section describes how to use Flink to write Hive tables, the definition of the Hive result table, parameters used for creating the result table, and sample code. For details, see <a href="https://nightlies.apache.org/flink/flink-docs-release-1.15/docs/connectors/table/hive/hive_read_write/" target="_blank" rel="noopener noreferrer">Apache Flink Hive Read & Write</a>.</p>
|
|
<p id="dli_08_15050__p105355289481">Flink supports writing data to Hive in both <strong id="dli_08_15050__b17405205520100">BATCH</strong> and <strong id="dli_08_15050__b1140635512107">STREAMING</strong> modes.</p>
|
|
<ul id="dli_08_15050__ul456011674911"><li id="dli_08_15050__li156014618499">When run as a BATCH application, Flink will write to a Hive table only making those records visible when the Job finishes. BATCH writes support both appending to and overwriting existing tables.</li><li id="dli_08_15050__li5560156124919"><strong id="dli_08_15050__b418693819127">STREAMING</strong> writes continuously adding new data to Hive, committing records - making them visible - incrementally. Users control when/how to trigger commits with several properties. Insert overwrite is not supported for streaming write. Please see the <a href="https://nightlies.apache.org/flink/flink-docs-release-1.15/docs/connectors/table/filesystem/" target="_blank" rel="noopener noreferrer">streaming sink</a> for a full list of available configurations.</li></ul>
|
|
</div>
|
|
<div class="section" id="dli_08_15050__dli_08_0256_en-us_topic_0132788972_section2579142713429"><h4 class="sectiontitle">Prerequisites</h4><div class="p" id="dli_08_15050__p12653145115206">To create a FileSystem source table, an enhanced datasource connection is required. You can set security group rules as required when you configure the connection.
|
|
</div>
|
|
</div>
|
|
<div class="section" id="dli_08_15050__section1230618441125"><h4 class="sectiontitle">Caveats</h4><ul id="dli_08_15050__ul181711754173415"><li id="dli_08_15050__li6647155118407">When you create a Flink OpenSource SQL job, set <strong id="dli_08_15050__b72247926422540">Flink Version</strong> to <strong id="dli_08_15050__b113263796422540">1.15</strong> in the <strong id="dli_08_15050__b111742231622540">Running Parameters</strong> tab. Select <strong id="dli_08_15050__b119714001822540">Save Job Log</strong>, and specify the OBS bucket for saving job logs.</li><li id="dli_08_15050__li118441048194615">For details about how to use data types, see <a href="dli_08_15014.html">Format</a>.</li><li id="dli_08_15050__li1877614152149">Flink 1.15 currently only supports creating OBS tables and DLI lakehouse tables using Hive syntax, which is supported by Hive dialect DDL statements.<ul id="dli_08_15050__ul52301539191612"><li id="dli_08_15050__li10367153891614">To create an OBS table using Hive syntax:<ul id="dli_08_15050__ul1176113556160"><li id="dli_08_15050__li11290555171618">For the default dialect, set <strong id="dli_08_15050__b89161947181618">hive.is-external</strong> to <strong id="dli_08_15050__b1391764761611">true</strong> in the with properties.</li><li id="dli_08_15050__li147011552191818">For the Hive dialect, use the <strong id="dli_08_15050__b1363619494160">EXTERNAL</strong> keyword in the create table statement.</li></ul>
|
|
</li><li id="dli_08_15050__li1083810474167">To create a DLI lakehouse table using Hive syntax:<ul id="dli_08_15050__ul1010114179193"><li id="dli_08_15050__li15304829161914">For the Hive dialect, add <strong id="dli_08_15050__b81713518155">'is_lakehouse'='true'</strong> to the table properties.</li></ul>
|
|
</li></ul>
|
|
</li><li id="dli_08_15050__li1488154619315">When creating a Flink OpenSource SQL job, enable checkpointing in the job editing interface.</li></ul>
|
|
</div>
|
|
<div class="section" id="dli_08_15050__section394917571396"><h4 class="sectiontitle">Syntax</h4><div class="codecoloring" codetype="Sql" id="dli_08_15050__screen1694920577399"><div class="highlight"><table class="highlighttable"><tr><td class="linenos"><div class="linenodiv"><pre><span class="normal"> 1</span>
|
|
<span class="normal"> 2</span>
|
|
<span class="normal"> 3</span>
|
|
<span class="normal"> 4</span>
|
|
<span class="normal"> 5</span>
|
|
<span class="normal"> 6</span>
|
|
<span class="normal"> 7</span>
|
|
<span class="normal"> 8</span>
|
|
<span class="normal"> 9</span>
|
|
<span class="normal">10</span>
|
|
<span class="normal">11</span>
|
|
<span class="normal">12</span>
|
|
<span class="normal">13</span>
|
|
<span class="normal">14</span>
|
|
<span class="normal">15</span>
|
|
<span class="normal">16</span>
|
|
<span class="normal">17</span>
|
|
<span class="normal">18</span>
|
|
<span class="normal">19</span>
|
|
<span class="normal">20</span>
|
|
<span class="normal">21</span>
|
|
<span class="normal">22</span>
|
|
<span class="normal">23</span>
|
|
<span class="normal">24</span>
|
|
<span class="normal">25</span>
|
|
<span class="normal">26</span>
|
|
<span class="normal">27</span>
|
|
<span class="normal">28</span>
|
|
<span class="normal">29</span>
|
|
<span class="normal">30</span>
|
|
<span class="normal">31</span></pre></div></td><td class="code"><div><pre><span></span><span class="k">CREATE</span><span class="w"> </span><span class="k">EXTERNAL</span><span class="w"> </span><span class="k">TABLE</span><span class="w"> </span><span class="p">[</span><span class="k">IF</span><span class="w"> </span><span class="k">NOT</span><span class="w"> </span><span class="k">EXISTS</span><span class="p">]</span><span class="w"> </span><span class="k">table_name</span>
|
|
<span class="w"> </span><span class="p">[(</span><span class="n">col_name</span><span class="w"> </span><span class="n">data_type</span><span class="w"> </span><span class="p">[</span><span class="n">column_constraint</span><span class="p">]</span><span class="w"> </span><span class="p">[</span><span class="k">COMMENT</span><span class="w"> </span><span class="n">col_comment</span><span class="p">],</span><span class="w"> </span><span class="p">...</span><span class="w"> </span><span class="p">[</span><span class="n">table_constraint</span><span class="p">])]</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="k">COMMENT</span><span class="w"> </span><span class="n">table_comment</span><span class="p">]</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="n">PARTITIONED</span><span class="w"> </span><span class="k">BY</span><span class="w"> </span><span class="p">(</span><span class="n">col_name</span><span class="w"> </span><span class="n">data_type</span><span class="w"> </span><span class="p">[</span><span class="k">COMMENT</span><span class="w"> </span><span class="n">col_comment</span><span class="p">],</span><span class="w"> </span><span class="p">...)]</span>
|
|
<span class="w"> </span><span class="p">[</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="k">ROW</span><span class="w"> </span><span class="n">FORMAT</span><span class="w"> </span><span class="n">row_format</span><span class="p">]</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="n">STORED</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="n">file_format</span><span class="p">]</span>
|
|
<span class="w"> </span><span class="p">]</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="k">LOCATION</span><span class="w"> </span><span class="n">obs_path</span><span class="p">]</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="n">TBLPROPERTIES</span><span class="w"> </span><span class="p">(</span><span class="n">property_name</span><span class="o">=</span><span class="n">property_value</span><span class="p">,</span><span class="w"> </span><span class="p">...)]</span>
|
|
|
|
<span class="n">row_format</span><span class="p">:</span>
|
|
<span class="w"> </span><span class="p">:</span><span class="w"> </span><span class="n">DELIMITED</span><span class="w"> </span><span class="p">[</span><span class="n">FIELDS</span><span class="w"> </span><span class="n">TERMINATED</span><span class="w"> </span><span class="k">BY</span><span class="w"> </span><span class="nb">char</span><span class="w"> </span><span class="p">[</span><span class="n">ESCAPED</span><span class="w"> </span><span class="k">BY</span><span class="w"> </span><span class="nb">char</span><span class="p">]]</span><span class="w"> </span><span class="p">[</span><span class="n">COLLECTION</span><span class="w"> </span><span class="n">ITEMS</span><span class="w"> </span><span class="n">TERMINATED</span><span class="w"> </span><span class="k">BY</span><span class="w"> </span><span class="nb">char</span><span class="p">]</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="k">MAP</span><span class="w"> </span><span class="n">KEYS</span><span class="w"> </span><span class="n">TERMINATED</span><span class="w"> </span><span class="k">BY</span><span class="w"> </span><span class="nb">char</span><span class="p">]</span><span class="w"> </span><span class="p">[</span><span class="n">LINES</span><span class="w"> </span><span class="n">TERMINATED</span><span class="w"> </span><span class="k">BY</span><span class="w"> </span><span class="nb">char</span><span class="p">]</span>
|
|
<span class="w"> </span><span class="p">[</span><span class="k">NULL</span><span class="w"> </span><span class="k">DEFINED</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nb">char</span><span class="p">]</span>
|
|
<span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">SERDE</span><span class="w"> </span><span class="n">serde_name</span><span class="w"> </span><span class="p">[</span><span class="k">WITH</span><span class="w"> </span><span class="n">SERDEPROPERTIES</span><span class="w"> </span><span class="p">(</span><span class="n">property_name</span><span class="o">=</span><span class="n">property_value</span><span class="p">,</span><span class="w"> </span><span class="p">...)]</span>
|
|
|
|
<span class="n">file_format</span><span class="p">:</span>
|
|
<span class="w"> </span><span class="p">:</span><span class="w"> </span><span class="n">SEQUENCEFILE</span>
|
|
<span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">TEXTFILE</span>
|
|
<span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">RCFILE</span>
|
|
<span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">ORC</span>
|
|
<span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">PARQUET</span>
|
|
<span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">AVRO</span>
|
|
<span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">INPUTFORMAT</span><span class="w"> </span><span class="n">input_format_classname</span><span class="w"> </span><span class="n">OUTPUTFORMAT</span><span class="w"> </span><span class="n">output_format_classname</span>
|
|
|
|
<span class="n">column_constraint</span><span class="p">:</span>
|
|
<span class="w"> </span><span class="p">:</span><span class="w"> </span><span class="k">NOT</span><span class="w"> </span><span class="k">NULL</span><span class="w"> </span><span class="p">[[</span><span class="n">ENABLE</span><span class="o">|</span><span class="n">DISABLE</span><span class="p">]</span><span class="w"> </span><span class="p">[</span><span class="n">VALIDATE</span><span class="o">|</span><span class="n">NOVALIDATE</span><span class="p">]</span><span class="w"> </span><span class="p">[</span><span class="n">RELY</span><span class="o">|</span><span class="n">NORELY</span><span class="p">]]</span>
|
|
|
|
<span class="n">table_constraint</span><span class="p">:</span>
|
|
<span class="w"> </span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="k">CONSTRAINT</span><span class="w"> </span><span class="k">constraint_name</span><span class="p">]</span><span class="w"> </span><span class="k">PRIMARY</span><span class="w"> </span><span class="k">KEY</span><span class="w"> </span><span class="p">(</span><span class="n">col_name</span><span class="p">,</span><span class="w"> </span><span class="p">...)</span><span class="w"> </span><span class="p">[[</span><span class="n">ENABLE</span><span class="o">|</span><span class="n">DISABLE</span><span class="p">]</span><span class="w"> </span><span class="p">[</span><span class="n">VALIDATE</span><span class="o">|</span><span class="n">NOVALIDATE</span><span class="p">]</span><span class="w"> </span><span class="p">[</span><span class="n">RELY</span><span class="o">|</span><span class="n">NORELY</span><span class="p">]]</span>
|
|
</pre></div></td></tr></table></div>
|
|
</div>
|
|
</div>
|
|
<div class="section" id="dli_08_15050__section14483145534319"><h4 class="sectiontitle">Parameter Description</h4><p id="dli_08_15050__p16483955184316">Please see the <a href="https://nightlies.apache.org/flink/flink-docs-release-1.15/docs/connectors/table/filesystem/" target="_blank" rel="noopener noreferrer">streaming sink</a> for a full list of available configurations.</p>
|
|
</div>
|
|
<div class="section" id="dli_08_15050__section1794965793911"><h4 class="sectiontitle">Example</h4><div class="p" id="dli_08_15050__p1596496145915">The following example demonstrates how to use Datagen to write to a Hive table with partition submission functionality.<pre class="screen" id="dli_08_15050__screen1290814215319">CREATE CATALOG myhive WITH (
|
|
'type' = 'hive' ,
|
|
'default-database' = 'demo',
|
|
'hive-conf-dir' = '/opt/flink/conf'
|
|
);
|
|
|
|
USE CATALOG myhive;
|
|
|
|
SET table.sql-dialect=hive;
|
|
|
|
-- drop table demo.student_hive_sink;
|
|
CREATE EXTERNAL TABLE IF NOT EXISTS demo.student_hive_sink(
|
|
name STRING,
|
|
score DOUBLE)
|
|
PARTITIONED BY (classNo INT)
|
|
STORED AS PARQUET
|
|
LOCATION 'obs://demo/spark.db/student_hive_sink'
|
|
TBLPROPERTIES (
|
|
'sink.partition-commit.policy.kind'='metastore,success-file'
|
|
);
|
|
|
|
SET table.sql-dialect=default;
|
|
create table if not exists student_datagen_source(
|
|
name STRING,
|
|
score DOUBLE,
|
|
classNo INT
|
|
) with (
|
|
'connector' = 'datagen',
|
|
'rows-per-second' = '1', --Generates a piece of data per second.
|
|
'fields.name.kind' = 'random', --Specifies a random generator for the user_id field.
|
|
'fields.name.length' = '7', --Limits the <strong id="dli_08_15050__b1662642154215">user_id</strong> length to <strong id="dli_08_15050__b5934172354211">7</strong>.
|
|
'fields.classNo.kind' ='random',
|
|
'fields.classNo.min' = '1',
|
|
'fields.classNo.max' = '10'
|
|
);
|
|
|
|
insert into student_hive_sink select * from student_datagen_source;</pre>
|
|
</div>
|
|
<p id="dli_08_15050__p264714223918">Query the result table using Spark SQL.</p>
|
|
<pre class="screen" id="dli_08_15050__screen3417115825619">select * from demo.student_hive_sink where classNo > 0 limit 10</pre>
|
|
</div>
|
|
</div>
|
|
<div>
|
|
<div class="familylinks">
|
|
<div class="parentlink"><strong>Parent topic:</strong> <a href="dli_08_15046.html">Hive</a></div>
|
|
</div>
|
|
</div>
|
|
|