Keyword Extractor
Extract recurring keywords and keyphrases from text using local frequency analysis.
This is a discovery tool for text you already have. It has no search volume, keyword difficulty, or competitor data, and it does not predict rankings. To check how often a keyword you already have in mind appears, use the Keyword Density Checker instead.
Text to analyse
Paste an article, page copy, notes, or research text. Counting is frequency-based and runs locally in your browser.
Stop words are common English function words such as the, of, is, and their. Filtering removes them from the single-word list only; inside a phrase they are always kept, so “cost of living” stays intact. The list is English-only. Group simple plurals merges a plural into its singular only when both forms appear in your text, so no invented word can reach the results.
Results
Every number below is a count taken from your text. Keywords and keyphrases are ranked separately, each by its own count.
Words analysed
165
Unique words
99
Keywords listed
20
Keyphrases listed
5
Keywords — single words
Ranked by how many times each word appears.
| Keyword | Count | Text share |
|---|---|---|
| 1.remote | 6 | 3.64% |
| 2.work | 6 | 3.64% |
| 3.cost | 4 | 2.42% |
| 4.pay | 4 | 2.42% |
| 5.how | 3 | 1.82% |
| 6.living | 3 | 1.82% |
| 7.anchor | 2 | 1.21% |
| 8.budget | 2 | 1.21% |
| 9.companies | 2 | 1.21% |
| 10.investment | 2 | 1.21% |
| 11.office | 2 | 1.21% |
| 12.policy | 2 | 1.21% |
| 13.Return | 2 | 1.21% |
| 14.salary | 2 | 1.21% |
| 15.say | 2 | 1.21% |
| 16.usually | 2 | 1.21% |
| 17.when | 2 | 1.21% |
| 18.about | 1 | 0.61% |
| 19.advice | 1 | 0.61% |
| 20.all | 1 | 0.61% |
Keyphrases — repeated word sequences
Only sequences that appear at least twice are listed.
| Keyphrase | Count | Text share |
|---|---|---|
| 1.remote work | 5 | 6.06% |
| 2.cost of living | 3 | 5.45% |
| 3.remote work policy | 2 | 3.64% |
| 4.Return on investment | 2 | 3.64% |
| 5.anchor pay | 2 | 2.42% |
How extraction works
The tool normalises spacing and quote characters, removes complete URLs and email addresses, then splits the remaining text into words. Hyphenated words, contractions, decimals such as 3.14, and technical terms such as Node.js, C++ and .NET are each kept as a single word rather than being broken apart.
Single words are counted directly. Two- and three-word sequences are counted separately, and only within a stretch of text uninterrupted by punctuation, so a phrase can never be assembled across a full stop or a comma.
Nothing here is statistical modelling or machine learning. Every number on the page is a count taken from the text you pasted, and the same text with the same options always produces the same output in the same order.
Keywords and keyphrases are ranked separately
Mixing single words and phrases into one ranked list makes the output misleading, because there is no honest way to compare a word that appears ten times with a three-word phrase that appears once. The two lists are kept apart and each is ranked by its own count.
Keyphrases must appear at least twice to be listed. In a short passage where nothing repeats, the tool falls back to phrases that appear once and says so above the list, so you are never left guessing whether a result recurred.
A two-word phrase is hidden when it only ever occurs inside a longer phrase with the same count, so “remote work policy” appearing three times does not also fill the list with “work policy”.
Reading counts and text share
Count is the number of times the term appears. Text share is the percentage of the analysed words that those occurrences cover: for a single word, count divided by total words; for a phrase, count multiplied by the phrase length, divided by total words.
Because a two-word phrase can sit inside a three-word phrase, phrase shares overlap and are not expected to add up to 100%. Treat share as a rough sense of how much of the page a term occupies, not as a budget to spend.
There is deliberately no combined relevance or optimisation score. A score built from frequency would only restate the count while implying it had measured something more.
Stop words and excluded terms
The stop-word filter removes common English function words — articles, pronouns, prepositions, conjunctions and auxiliary verbs such as the, of, is and their — from the single-word list. It is English-only; word splitting handles other languages and scripts, but the stop-word list does not.
Words that often carry the actual topic of a page are deliberately kept, including how, what, why, where, when, which, who, more, most, about, not, make, use, find and help. Inside a phrase, function words are always kept, which is why “cost of living” and “how to calculate” survive intact.
Use the exclude field for terms you already know about and do not want to see: a company name, an author byline, a product name, or boilerplate that repeats in every draft. Excluded terms are dropped from both lists, and any phrase containing one is dropped too.
Privacy and limits
The analysis runs in your browser on this page. No text is uploaded, no account is needed, and copying or exporting results creates a file locally on your device.
The honest limits: frequency is not importance, this tool has no search or competitor data of any kind, it cannot tell synonyms or related concepts apart, and grouping only merges a plural with its singular when both forms already appear in your text — so it will never invent a word, but it will also leave many variants ungrouped.
Frequently Asked Questions
Related tools