Lead Scoring Template: Why Your Best Leads Reply Worst covers the 100-point model in 2 minutes, and the back-test that inverted our own score. Watch on YouTube

TL;DR

→ We scored 59,339 Upwork opportunities with a match model, then measured what replied. The top-scoring decile came back at 5.18% against 8.19% for the bottom.

→ Every lead scoring template on the internet has two dimensions: fit and engagement. Neither one measures how many other vendors are chasing the same lead.

→ Contention is the biggest measured effect we have. Reply rate falls from 9.44% to 2.11% when eleven or more GigRadar customer agencies bid the same job.

→ The firmographic default is inverted here. Clients who have spent over $500k on Upwork reply at 3.85%, against 8.15% for clients who have spent under $1k.

→ Below: a free interactive lead scoring template (25 points fit, 40 client, 35 contention), the point values with sample sizes, and a copyable rubric for your sheet or CRM.

Two years ago we shipped a match score. It ranks every Upwork job against an agency's scanner and hands back a number between 0 and 1.

Then we joined 59,339 of those scored opportunities to what happened next. The jobs our model called a near-perfect match replied at 5.18%, and the ones it nearly rejected replied at 8.19%.

The score was not broken. It was measuring the wrong quantity: a job that looks perfect to your scoring model looks perfect to everyone else's too, and the number you built to find opportunity turned out to be a very good detector of crowding.

That is the flaw in every lead scoring template I have ever been handed, and it is why the one below spends more than a third of its points on a dimension your CRM does not have a field for.

Every lead scoring template measures two things, and neither one is competition

Open any lead scoring tool and you get the same two axes. HubSpot's own documentation is explicit about it: a fit score rates "property values" and an engagement score rates "events", and a combined score is "the total value for the engagement and fit points combined".

Fit describes the lead. Engagement describes what the lead did.

Both are properties of a single record sitting alone in your database.

Neither answers the only question that decides whether you get a reply on a marketplace: how many other agencies are working this exact lead this afternoon.

The marketplace correction

In a CRM, two salespeople scoring the same inbound form fill is a routing bug. On Upwork it is Tuesday.

The lead is public, the buyer is one inbox, and your score competes with forty other scores computed from the same visible fields.

Diagram comparing a standard two-axis lead scoring model (fit and engagement) with a marketplace model that adds a contention block worth 35 points

The missing third axis. Fit and engagement both describe the lead in isolation; contention describes the queue in front of it.

That difference is not cosmetic. It flips the sign on half the criteria a standard template ships with.

Upwork publishes the buyer's history on every job post under "About the client". That panel is the raw input for the fit and client blocks below.

Upwork job post "About the client" sidebar showing total spent, jobs posted, hire rate, average hourly rate paid, and member-since date

Every input in Block B is on this panel: total spent, jobs posted, hire rate, and average rate paid to past hires.

What happened when we scored 59,339 leads and then checked

Line the score deciles up against reply rate and the curve runs the wrong way.

Reply rate by match-score decile

n = 59,339 proposals joined to opportunity records, Dec 2025 to Feb 2026. Baseline 7.45%.

Bottom (0.00–0.38)8.19%
0.38–0.458.49%
0.45–0.507.28%
0.50–0.536.48%
0.53–0.607.96%
0.60–0.697.18%
0.69–0.796.69%
0.79–0.837.12%
Top (0.83–0.93)5.18%

Source: GigRadar internal pipeline data, Dec 2025 to Feb 2026.

The bottom decile beats the top decile by 58%. The relationship is noisy in the middle and unambiguous at the ends.

I want to be precise about what this does and does not prove. It does not prove your fit criteria are worthless, and it does not prove low-fit work converts to revenue better.

It proves that a fit score, on a public marketplace, is partly a measurement of how many other people can also see the opportunity.

For scale: a 30-feature model built on cover-letter and bid features reached an AUC of 0.585 and explained roughly 16% of the variance in reply outcome. That model contained no client-quality feature and no contention feature at all, which is most of what the template below scores.

So the template below is a triage tool, not an oracle. It is built to stop you spending Connects on the most contested third of your pipeline.

The lead scoring template, as a calculator

Score one live job post. The three blocks are weighted 25 fit, 40 client, 35 contention, and every point value traces back to a measured reply-rate delta in the tables further down.

Free lead scoring template

Answer nine questions about one job post. Nothing is sent anywhere.

Block A: Fit (max 25)

Block B: Client (max 40)

Block C: Contention (max 35, goes negative)

Thresholds: 65 and up means bid and consider boosting, 45 to 64 means bid without boosting, 25 to 44 means bid only with spare Connects, under 25 means skip. Derived from single-variable reply deltas, so treat it as triage rather than prediction.

Why the fit block is capped at 25 points

Standard templates give fit the majority of the weight, because in a CRM fit is the only thing you know before the lead does anything. The tooling encourages it: HubSpot lets you stretch the score range as far as -10,000 to 10,000.

That is a lot of room to express a preference you have never tested. On Upwork you also know the buyer's entire purchase history, so spending 70 points on "do they look like our ICP" wastes the better signal.

Fit still earns its 25. Category experience is real, and the job title carries surprising information: titles containing the word "urgent" reply 3.07 percentage points above baseline, while titles built on a bare role noun like "developer" run 2.70 points below it.

The one fit signal worth more than it looks is the experience-level tag, and it points the opposite way from where agencies bid.

Experience level on the postnReply rate
Entry level33814.20%
Entry (older tag)5939.44%
Expert63,3237.78%
Intermediate69,5767.10%

Source: GigRadar internal pipeline data, Dec 2025 to Feb 2026. The entry-level cohort is under 1% of volume, so treat the 14.20% as directional.

Almost every agency filters entry-level work out on day one. That filter is why the cohort is uncrowded, and the uncrowding is most of the lift.

The client block is where the standard template is simply inverted

Generic lead scoring hands the most points to the biggest company. Enterprise headcount, big revenue band, well-known logo, plus twenty.

We have the buyer's actual lifetime spend on 91,056 proposals, which is the marketplace version of that field. It runs the other way.

Client lifetime spendnReply ratePoints
$0, new client25,4136.89%+7
$1 to $1k11,5068.15%+12
$1k to $5k13,8177.90%+12
$5k to $25k16,9086.84%+6
$25k to $100k12,9776.03%+2
$100k to $500k8,0705.70%0
Over $500k2,3653.85%-8

Source: GigRadar internal pipeline data, Dec 2025 to Feb 2026. Client metadata is the job's posting-time snapshot.

The buyers who have spent over half a million dollars on the platform reply at less than half the rate of the ones who have spent under a thousand. They are professional buyers with incumbent freelancers, and a lot of their public posts are price discovery rather than hiring.

Feedback score behaves the same way, and this is the point where most people stop trusting the data.

13.11%

Reply rate on clients rated under 3.5 stars (n=351)

6.51%

Reply rate on clients rated 4.8 or higher (n=49,135)

5.83%

Reply rate on clients who have posted exactly one job before (n=10,806)

A client with poor reviews is reading every proposal, because their last hire went badly and their reputation on the platform is a problem they need solved. A client with 4.8 stars and a full roster is skimming.

This is the one criterion in the template I would tell you to override on purpose. Reply rate is not delivery risk.

Upwork's payment protection covers the money on hourly contracts. It does not cover three weeks of scope arguments.

If your agency cannot afford a difficult client, score it down and accept the lost replies. Our ICP red-flag list covers which of those clients are worth avoiding.

Duration is the last client signal, and it is bimodal. Under a month replies at 8.33% and over six months at 8.08%, while the three to six month band sits at 6.71%, which is where scope-creeping projects with no real deadline live.

Contention is the block your CRM has no field for

This is the largest effect in the entire dataset, and no lead scoring template I have seen contains it.

We can see when several of our customers bid the same Upwork job. When we bin proposals by how many agencies showed up, reply rate collapses.

Reply rate by number of GigRadar agencies bidding the same job

n = 59,339 proposals across 30,964 distinct Upwork jobs.

1 agency9.44%
2 agencies7.41%
3 agencies6.57%
4 to 55.08%
6 to 104.50%
11 or more2.11%

Source: GigRadar internal pipeline data, Dec 2025 to Feb 2026.

To be exact about what this counts: these bins are GigRadar customer agencies, not total proposals on the job. It is a floor on real contention, because every job also carries bidders we cannot see.

Reply rate falls 4.5 times between one and eleven. And 67% of proposals in that window landed on a job where at least one other GigRadar agency was also bidding.

The damage is not evenly spread. Technical categories cannibalise hardest, because their scanner queries overlap most.

Category1 agency bidding6 to 10 biddingDrop
IT & Networking12.67%3.85%-8.82pp
Data Science & Analytics9.63%0.99%-8.64pp
Design & Creative8.97%1.69%-7.28pp
Web, Mobile & SW Dev7.97%3.90%-4.07pp
Sales & Marketing11.44%7.77%-3.67pp

Source: GigRadar internal pipeline data, Dec 2025 to Feb 2026.

A Data Science lead that six GigRadar agencies are working is worth 1% of a reply. The same lead with nobody else on it is worth nearly ten times that, and your fit score cannot tell the two apart because the job post is identical.

Freelancers have been hand-approximating this for years without a name for it. The threads are full of people refusing to bid once the proposal counter crosses a threshold, or once a client's hire rate drops below their personal cut-off.

Reddit r/Upwork thread on a client with 953 job posts and a 13% hire rate, and why freelancers stop applying to low-hire-rate clients

A client with 953 posts and a 13% hire rate, and the reply explaining the feedback loop. Source: r/Upwork.

Ask that subreddit what hire rate is high enough to justify applying and you get answers ranging from 50% to "nothing under 90% if you are new". Those are hand-built scoring thresholds, invented independently by people who have never used the phrase lead scoring template.

Upwork already prices contention, and the price band is a scoring signal

You do not have to guess at crowding, because Upwork sells you a reading of it. The number of Connects a job costs is set by demand: Upwork's own help page says the cost "varies per job and can change during the time the job is posted" based on "project size, scope, and market demand".

Connects cost $0.15 each, so that price is also your unit cost per scored lead. What our data shows is that the price ladder has a hole in the middle.

Connects to applynReply ratePoints
1 to 291,3887.61%+6
5 to 63185.35%-8
7 to 81,0184.03%-8
9 to 103,0675.05%-8
11 to 123,1858.04%+8
13 to 1610,6027.58%+8
17 or more24,2947.16%+8

Source: GigRadar internal pipeline data, Dec 2025 to Feb 2026.

Cheap jobs work and expensive jobs work. The five to ten Connect band is where reply rate bottoms out at 4.03%, and it costs you three to five times what the cheap band does.

The other lever is the boost auction, a sealed-bid contest for the top four slots that Upwork describes as useful "especially when lots of freelancers are applying to the same job".

If you are bidding into that auction, the platform has already told you the lead is contested. The full economics sit in the Connects cost-per-hire calculator.

The negative block is where the real gain is

Most scoring rubrics treat negative points as an afterthought for unsubscribes and competitor domains. Here they carry the model.

The clearest one is pricing position. We have the client's average paid rate to past hires on 56,643 proposals, and the worst possible bid is the one that sits just above it.

Your bid ÷ client's typical paid ratenReply rate
Under 0.5x1,6486.74%
0.5x to 0.8x3,6756.64%
0.8x to 1.0x3,3336.36%
1.0x to 1.2x3,5605.37%
1.2x to 1.5x4,7545.68%
1.5x to 2x6,0376.53%
2x to 5x12,4926.89%
Over 5x4,5837.70%

Source: GigRadar internal pipeline data, Dec 2025 to Feb 2026.

Slightly expensive is the one position with no story attached to it. Clearly cheap and clearly premium both convert better than the band that says "a bit more than you usually pay, for reasons I have not explained".

The second negative worth wiring in is your own inbox position. Being in the first five proposals replies at 7.28%, positions 6 to 10 sink to 6.03%, and positions 21 to 50 recover to 7.77%.

The second-screen trap

Clients read the first five proposals properly, skim the next fifteen, then sample randomly into the deep pile if nobody in the top twenty grabbed them. Landing at position 8 is worse than landing at position 30, so if you cannot be early, being late is safer than being nearly early.

GigRadar

Free for Upwork agencies

Score every job before you spend a Connect on it

GigRadar scans Upwork for the jobs that match your agency, scores them on client quality and contention, and submits proposals through our own Upwork Business Manager account under human supervision. Your freelancer account is never touched.

Get Your Free Agency Audit →

The full rubric, ready to paste into a sheet

Here is the whole lead scoring template as text. Paste it into a spreadsheet column, or translate each line into a scoring rule in your CRM.

LEAD SCORING TEMPLATE / UPWORK AGENCY EDITION Point values derived from reply-rate deltas across GigRadar pipeline samples of 56,643 to 133,872 proposals, Dec 2025 - Feb 2026. Baseline reply rate 7.45%. BLOCK A: FIT (max 25) +10 Hired in this exact category before +5 Adjacent category +7 Urgency language in the job title +3 Specific outcome or deliverable in the title -5 Generic role noun in the title (developer, designer, assistant) +8 Post tagged Entry level +4 Post tagged Expert +2 Post tagged Intermediate BLOCK B: CLIENT (max 40) +12 Lifetime spend $1-$5k +7 Lifetime spend $0 (new client) +6 Lifetime spend $5k-$25k +2 Lifetime spend $25k-$100k -8 Lifetime spend over $500k +10 Client feedback under 4.0 +6 Client feedback 4.0-4.5 +4 Client feedback 4.5-4.8 +3 No client rating yet 0 Client feedback 4.8+ +8 Client has posted 6-25 jobs +6 Client has posted none +4 Client has posted 2-5 jobs +4 Client has posted more than 50 jobs +3 Client has posted 26-50 jobs -6 Client has posted exactly one +10 Duration under 1 month +9 Duration over 6 months +3 Duration 1-3 months -4 Duration 3-6 months BLOCK C: CONTENTION (max 35, goes negative) +10 You would land 21st-50th in the client's pile +8 You would land 1st-5th +8 You would land 51st or later -4 You would land 11th-20th -8 You would land 6th-10th +10 Your own scanner rates this a LOOSE match +4 Your own scanner rates this a middling match -8 Your own scanner rates this a NEAR-PERFECT match +8 Upwork charging 11+ Connects to apply +6 Upwork charging 1-2 Connects -8 Upwork charging 5-10 Connects +7 Your bid over 5x or under 0.8x the client's typical paid rate +3 Your bid 1.5x-2x theirs -7 Your bid 1.0x-1.2x theirs THRESHOLDS 65+ Bid now, consider boosting 45-64 Bid, do not boost 25-44 Bid only with spare Connects Under 25 Skip RECALIBRATE EVERY 90 DAYS. Reply data needs 75-90 days to cure.

The worksheet version

If you would rather score in a sheet than in the widget, this is the same model as a grid. Copy it straight into Google Sheets or Excel with the button, then fill the last column per lead.

BlockCriterionScaleMaxYour score
A. FitCategory you have been hired inexact +10 / adjacent +5 / new 010
A. FitJob title language"urgent" +7 / outcome +3 / neutral 0 / role noun -57
A. FitExperience level tagentry +8 / expert +4 / intermediate +28
B. ClientLifetime spend on Upwork$1-5k +12 / new +7 / $5-25k +6 / $25-100k +2 / $500k+ -812
B. ClientFeedback score<4.0 +10 / 4.0-4.5 +6 / 4.5-4.8 +4 / none +3 / 4.8+ 010
B. ClientJobs posted before6-25 +8 / none +6 / 2-5 +4 / 50+ +4 / 26-50 +3 / exactly one -68
B. ClientStated duration<1mo +10 / >6mo +9 / 1-3mo +3 / 3-6mo -410
C. ContentionYour position in the pile21-50th +10 / 1-5th +8 / 51st+ +8 / 11-20th -4 / 6-10th -810
C. ContentionYour scanner's own match scoreloose +10 / middling +4 / near-perfect -810
C. ContentionConnects charged to apply11+  +8 / 1-2 +6 / 5-10 -88
C. ContentionBid vs client's typical paid rate>5x or <0.8x +7 / 1.5-2x +3 / 1.0-1.2x -77
TotalThreshold65+ bid and boost / 45-64 bid / 25-44 spare only / <25 skip100

How to run this without turning it into a spreadsheet ritual

A scoring model that nobody maintains decays into superstition within a quarter. Five steps, and the last one is the only one people skip.

1

Score 50 jobs you already bid on, blind to the outcome

Use last quarter's proposals. You need the score and the result in the same sheet before you trust a single threshold.

2

Move your own thresholds, not the point values

The point values come from a large sample. Your cut-off is the part that should be yours, and it depends on how many Connects you can afford to lose per week.

3

Wire the contention block into your scanner, not your proposal step

Where you would land in the pile and the Connect price are both visible before you write anything. Scoring after the cover letter is written wastes the only cost the model was built to save, and Boolean operators in Upwork's advanced search are what let you build a query nobody else is running.

4

Track cost per reply by score band, not overall reply rate

Cost per reply spans 2.4 times across categories in our data, from $14.30 in Writing to $34.21 in Web, Mobile and Software Development. A blended number hides which band is burning your budget.

5

Recalibrate on a 90-day clock

Reply rates take 75 to 90 days to fully cure, so a 30-day read on a new threshold is mostly noise. Put the review in the calendar or it will not happen.

The scanner point matters more than it sounds. If your queries surface the same jobs every other agency's queries surface, you are paying full price for the most contested third of the market, which is the failure mode our 60-second job qualification guide was written to catch by hand.

Everything upstream of this template is targeting. If your ICP template and your scanner disagree, the score will faithfully rank leads you should never have surfaced, and the MQL to SQL handoff will inherit the error.

What this template deliberately does not do

It scores the client, not you. Your own Job Success Score is the other half of whether a contested lead is winnable, and no rubric you build about the buyer will fix a profile that clients bounce off.

It also scores replies, not revenue. Reply rate is our north-star because hires under-count badly when customers close off-platform, but the two come apart in specific places.

What the score optimisesWhat it will not tell you
Whether a client opens a conversationWhether that client pays on time
Which leads are cheapest to reachWhich leads have the largest contract value
Where your Connects are being wastedWhether your cover letter is any good
Crowding at the moment you scoredCrowding four hours later

The last row is the real limitation. Contention is a live quantity, and a score computed at 9am on a job that goes viral by lunchtime is stale by the time your proposal lands.

That is the part a spreadsheet cannot fix, and it is the honest reason we built the scoring into the scanner rather than shipping you a better sheet. If you want the same logic running continuously against every new Upwork post in your niche, GigRadar's scoring system is that idea productised, and our response-rate benchmarks give you the numbers to check it against.

Lead scoring template FAQ

What should a lead scoring template include?

Three blocks rather than the usual two: fit (capped, because on a public marketplace a high fit score doubles as a crowding warning), client quality read from the buyer's visible purchase history, and contention, meaning how many other vendors are working the same lead. Contention carries the largest measured effect in our pipeline data, with reply rate falling from 9.44% to 2.11% between one GigRadar customer agency bidding a job and eleven of them.

What point values should I assign in a lead scoring model?

Assign them from measured outcome deltas rather than intuition: the template above weights 25 points to fit, 40 to client quality and 35 to contention, and every line traces to a reply-rate difference with its sample size published. If you have your own outcome data, score 50 past leads blind and let your own deltas set the weights.

What is a good lead score threshold?

We use 65 and above for bid-and-boost, 45 to 64 for bid without boosting, 25 to 44 for spare capacity only, and skip below 25. Move the threshold rather than the point values, because the right cut-off depends entirely on how many Connects you can afford to lose per week.

Why do my highest-scoring leads convert worst?

Because on a marketplace, the criteria that make a lead look attractive to you make it look attractive to everyone else running a similar model. Across 59,339 scored opportunities our top match-score decile replied at 5.18% against 8.19% for the bottom decile, and reply rate falls from 9.44% to 2.11% between one GigRadar customer agency on a job and eleven of them, which is the signature of crowding rather than a broken score.

Should big clients score higher in a lead scoring template?

Not on Upwork, where the generic firmographic rule that awards points for company size is inverted. Buyers with over $500k of lifetime platform spend replied at 3.85% against 8.15% for buyers who had spent under $1k, because the large accounts already have incumbent freelancers and treat the public bid pool as price discovery.

How often should I recalibrate a lead scoring model?

Every 90 days. Reply data needs 75 to 90 days to fully cure in our windows, so judging a new threshold after 30 days mostly measures noise, and leaving it untouched for a year means scoring against a market that has moved.