CSV Column Extractor
Select CSV columns in an explicit order while retaining exact text and blank records
Data is processed in this browser without saving input drafts. Limits: 10,000 data rows, 200 columns, 200,000 cells; 32,767 UTF-16 code units per cell.
Pasted line endings follow textarea rules; upload UTF-8 to retain original carriage returns. All CSV values remain text, including leading zeros, header whitespace, repeated headers and physical blank records. Nonblank records must have equal width. Choose a delimiter explicitly when detection is ambiguous.
The full guide also includes pitfalls, worked examples, snippets, FAQs, and related tools for checking results or troubleshooting.
About this tool
CSV Column Extractor selects columns in the order you specify while keeping every value as text. Use 1-based positions such as 3,1, or switch to exact header names and enter a JSON array such as ["email","10","last,name"]. Header names are not trimmed or converted into object keys, so numeric names, blank labels and __proto__ retain their original positions. Duplicate names require positional selection, and each column may be selected once. Comma, semicolon, Tab and pipe are supported; automatic detection chooses a consistent multi-column candidate, so explicitly select the delimiter for ambiguous input. Quoted delimiters, doubled quotes and embedded CR, LF or CRLF stay inside their original fields. Physical blank records remain separate blank records; other records must have equal width. Limits are 10 MiB of UTF-8 input, 10,000 data rows, 200 columns, 200,000 cells and 32,767 UTF-16 code units per cell. Copy and download contain the complete extracted CSV up to 16 MiB. JSON is only a preview of the first 20 data rows, represented as ordered columns and rows arrays; each displayed text preview stops at 64 KiB. Processing is local, and input drafts are not saved. Opening exported text in spreadsheet software can still trigger type or formula interpretation, so import untrusted fields as text. Browser textareas normalize pasted line endings; upload a UTF-8 file when original carriage returns must be retained.
Scenario Recipes
Extract an identifier and a quoted customer label
Goal: Keep the source ID spelling while placing the label first.
- Paste id,"last,name",status followed by 00123,"Doe, Jane",active. Select comma and keep the first-record header option enabled.
- Select exact header names and enter ["last,name","id"], then extract. Confirm that the first output field is Doe, Jane and the second is 00123.
- Copy full CSV or download it. Read it as CSV with a comma delimiter and text columns; the displayed JSON is only the first 20 rows.
Result: A two-column CSV in the requested order, with the comma inside the name quoted and the leading-zero ID unchanged.
Distinguish physical blank records from empty fields
Goal: Retain the record shape during a one-column extract.
- Use a one-column header h, then a physical blank line, a line containing "", and a line containing one space.
- Extract position 1. The preview records are [], [""] and [" "]; each represents a different source record.
- Read the full CSV using a standards-aware CSV reader. The physical blank line remains blank and the quoted empty value remains one field.
Result: Empty records remain present rather than being silently filtered out.
Production Snippets
Keep the third column first without object-key collisions
text
Input:
10,2,__proto__,note
00123,12.5,own,"line 1
line 2"
Position selection: 3,1
Output CSV:
__proto__,10
own,00123
Exact-name mode uses a JSON array: ["__proto__","10"]
Repeated headers must be selected by position.
A JSON preview contains columns metadata and rows arrays, not keyed objects.Failure Clinic (Common Pitfalls)
A header name is rejected or selects an ambiguous column
Cause: Exact-name mode uses JSON strings, and a repeated label refers to more than one position.
Fix: Use ["email","10"] for exact labels or switch to 1-based positions for duplicates. Preserve any intended leading/trailing header spaces. Do not mix names and positions in one selection.
The output appears to have fewer rows in JSON
Cause: JSON is explicitly a preview of the first 20 rows, and long text previews stop at 64 KiB.
Fix: Use full CSV copy/download for every accepted record. The total row count is included beside the preview metadata.
A malformed record stops extraction
Cause: Bare quotes, unclosed quotes, trailing text after closing quotes or unequal nonblank widths cannot be interpreted without changing data.
Fix: Quote the entire field, double internal quotes and fix the record width. Choose the actual delimiter; a field can contain commas even when the file is semicolon-separated.
Spreadsheet software changes an exported ID or runs formula-like text
Cause: CSV stores field text without a persistent schema, and the destination application may infer numbers or formulas.
Fix: Import as text where source spelling matters. The extractor preserves formula-like prefixes and does not silently add apostrophes or change the data.
Frequently Asked Questions
How do I select a header containing a comma or a numeric name?
Choose exact header names and enter a JSON string array, for example ["10","last,name"]. Names match exactly, including spaces. Missing or duplicated names are rejected; use positions for duplicates.
Does 3,1 change the output order?
Yes. The third source column is exported first, followed by the first. Blank tokens, duplicate selections, zero, fractions and leading-zero positions such as 01 are rejected.
Are empty records and embedded line breaks kept?
Yes. A physical blank record remains a blank line, while one quoted empty field is retained as one empty field. CR, LF and CRLF inside quoted cells are preserved; output record separators use CRLF.
Is JSON preview a full JSON conversion?
No. It contains ordered column metadata and only the first 20 data rows. Use full CSV copy/download for all records, or a dedicated converter for a complete JSON dataset.
Which sizes and encodings are supported?
Upload valid UTF-8 CSV, TSV or TXT up to 10 MiB, or paste text. Limits: 10,000 data rows, 200 columns, 200,000 cells, 32,767 UTF-16 code units per cell and 16 MiB exported CSV.
Does it save my CSV or change number spelling?
No input draft is saved. Processing stays in the browser and values remain text, including 00123 and long numbers. A later spreadsheet import may infer types, so choose text import when spelling matters.
Keep browsing