Skip to content
UniKit

Regex generator from examples

Infer a regex from examples: give the samples that must match and the ones that must not, and a set of built-in patterns (email, URL, IP, phone, date, time, number, Chinese, hex, ID card, postal code, custom character classes) is filtered against them — with per-sample results, and an honest “no pattern fits” when nothing does.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Samples

Candidates are tested with search semantics, not whole-string matching. Hits are ranked by specificity: a pattern only counts when it matches every positive and no negative.

Custom character classes (optional)

Pick the character sets and length you need; the resulting whole-string pattern joins the candidate list.

Character sets
Generated pattern^[a-z0-9]{6,}$
Ignore case

Result

Patterns that fit

Email address/[A-Za-z0-9._%+-]+@[A-Za-z0-9-]+(?:\.[A-Za-z0-9-]+)+/u

The usual user@host.tld shape

  • ✓ alice@example.com

Specificity: 95

Identifier / English word/[A-Za-z_][A-Za-z0-9_]*/u

Letters or underscore followed by letters and digits

  • ✓ alice@example.com

Specificity: 10

Close but rejected

Custom character classes/^[a-z0-9]{6,}$/u
  • ✗ alice@example.com · Missed positives
URL/https?://[^\s"'<>]+/u
  • ✗ alice@example.com · Missed positives
UUID/[0-9A-Fa-f]{8}-(?:[0-9A-Fa-f]{4}-){3}[0-9A-Fa-f]{12}/u
  • ✗ alice@example.com · Missed positives
IPv4 address/(?:(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)\.){3}(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)/u
  • ✗ alice@example.com · Missed positives
IPv6 address/(?:[0-9A-Fa-f]{1,4}:){7}[0-9A-Fa-f]{1,4}/u
  • ✗ alice@example.com · Missed positives
MAC address/(?:[0-9A-Fa-f]{2}[:-]){5}[0-9A-Fa-f]{2}/u
  • ✗ alice@example.com · Missed positives

What this tool does

  • You have a few “it should look like this” samples but cannot write the pattern: paste them in and let the built-in candidates do the filtering.
  • Separate “must match” from “must not match” with counter-examples — \d+ happily matches a phone number, but adding a negative sample stops it from being recommended.
  • Sanity-check a validation rule before you commit to it: give a few valid and a few invalid values and see whether one built-in pattern satisfies both.
  • Grab a regex for a common format (email, URL, IP, ID card, postal code, date/time, colour, UUID…) and see exactly what it matches and where its edges are.

Example

Input

Must match: alice@example.com
Must not match: not-an-email

Output

Hit: Email address /[A-Za-z0-9._%+-]+@[A-Za-z0-9-]+(?:\.[A-Za-z0-9-]+)+/u (✓ alice@example.com). The broader “Identifier” pattern matches the positive too, but it also matches not-an-email, so it is only listed as a rejected candidate.

Matching uses search semantics rather than whole-string matching, and hits are ranked by specificity — the email pattern comes before the identifier one.

Frequently asked questions

Why infer from examples instead of generating from a sentence?

The candidate table is finite, and the tool only does one thing: it verifies which patterns really satisfy your positives and negatives. When none do, it says so plainly instead of producing a regex that looks right but matches the wrong things. To check a pattern you wrote yourself, the regex tester is the better tool.

Do the hits match whole strings?

No — the built-in candidates use search semantics, so \d+ matches the 123 inside abc123. For whole-string validation (say “6–12 lowercase letters or digits”) use the custom character classes block: it emits an anchored pattern such as ^[a-z0-9]{6,12}$.

Why do several patterns match the same sample?

Because they all satisfy your constraints — they just differ in strictness. The email pattern is far more specific than the identifier pattern, so it is ranked first. Add a few negative samples to weed out the loose ones; that is usually more effective than adding more positives.

What are the custom character classes for?

When no built-in pattern covers “fixed length plus a specific character set”, pick the sets (upper, lower, digits, underscore, Chinese, whitespace, hyphen) and the length range; the resulting anchored pattern joins the candidate list. The hyphen is always placed last so a class like [a-z-0-9] never turns into an accidental range.

Is there a limit on samples?

At most 100 samples, each up to 200 characters; blank lines are ignored. Going over the limit is reported rather than silently truncated, because dropping a few samples can completely change which pattern fits.

Keywords:regex generatorregex from examplespattern generatorregular expressionsample based regex正则生成正则表达式生成反推正则正则表达式示例生成正则

Related tools