FTP
Pipeline Concepts
Section titled βPipeline ConceptsβBefore setting up the Pipeline, learn about Pipeline concepts here
Step by Step Guide
Section titled βStep by Step GuideβSTEP-1: Configure FTP/FTPs Connection
Section titled βSTEP-1: Configure FTP/FTPs ConnectionβTo learn about Connection, refer here
- Log into Sprinkle application
- Navigate to Ingest -> Connections Tab -> New Connection ->
- Select FTP or FTPs
- Provide all the mandatory details
- Name: Name to identify this connection
- Host: FTP Hostname
- Port: FTP port. Default is 22.
- User: FTP username
- Password
- Test Connection
- Create
STEP-2: Configure FTP/FTPs Pipeline
Section titled βSTEP-2: Configure FTP/FTPs PipelineβTo learn about Pipeline, refer here
- Navigate to Ingest -> Pipeline Tab -> Add ->
- Select FTP or FTPs
- Provide the name -> Create
- Connection Tab:
- From the drop-down, select the name of connection created in STEP-2
- Update
STEP-3: Create Dataset
Section titled βSTEP-3: Create DatasetβDatasets Tab: To learn about Dataset, refer here. Add Dataset for each directory that you want to replicate, providing following details
- Table Name (Required) : Table name suffix which will be used to create the table in the warehouse
- Directory Path (Required) :Provide the full path
- Ingestion Mode (Required) :
- Complete: Full folder is downloaded and ingested in every ingestion job run
- Incremental: Ingest only the new files in every ingestion job run. Use this option if your folder is very large, and you are getting new files continuously
- Remove Duplicate Rows:
- Unique Key: Unique key from table, to dedup data across multiple ingestions
- Time Column Name: Will be used to order data for deduping
- Max Job Runtime: Give maximum time in minutes for which data should be downloaded. Ingestion job will run specified max minutes and checkpoint will be updated. Next run will continue from checkpoint.
- Remove Duplicate Rows:
- File Type: Select the File Format
- JSON
- CSV
- Select Delimiter - Comma, Tab, Pipe, Dash, Other Character
- Parquet
- ORC
- Destination Schema (Required) : Data warehouse schema where the table will be ingested into
- Warehouse Table name (Optional) : It is optional field. If not given, sprinkle will create like ds_<Pipelinename>_<tablename>
- Destination Create Table Clause: Provide additional clauses to warehouse-create table queries such as clustering, partitioning, and more, useful for optimizing DML statements. Learn more on how to use this field.
- Create
STEP-4: Run and schedule Ingestion
Section titled βSTEP-4: Run and schedule IngestionβIn the Ingestion Jobs tab:
- Trigger the Job, using Run button
- To schedule, enable Auto-Run. Change the frequency if needed