標準評価は「誰が審査するか」で変わるべきか
——MR中途採用一次選考におけるSEGT実証
評価者が変われば、標準評価はどう変わるのか。医薬品企業のMR中途採用の一次書類選考を題材に、eval000のSEGT(Standard Evaluation Generation Theory)で検証しました。
採用選考の一次書類審査は、多くの場合「複数の評価者が採点し、平均する」という方法で標準化されています。しかし、評価者が変われば重視する観点も変わります。即戦力を求める採用マネージャーと、将来性を見る育成担当者とでは、同じ応募書類でも評価が変わって当然です。
この「評価者による違い」を、単なるブレとして平均化してしまってよいのでしょうか。それとも、会社が求める基準と評価者の視点の違いを、両方とも活かした形で評価を導き出すべきなのでしょうか。今回、SEGTを用いてこの問いを検証しました。
① 外生原理からNorm Nを生成する——「原理→ルーブリックR→生成評価値N」という設計
SEGTの理論構造を正確に述べると、採用方針の文章そのものが「Norm N」なのではありません。理論上、この関係は次の4段階として整理されます。
「即戦力となるMR職を採用する」といった採用方針そのもの(外生原理)は、まずルーブリックR(評価軸とアンカー基準)へと翻訳されます。Norm Nは、この外生原理そのものではなく、外生原理をルーブリックRに照らして導出された、生成評価値という位置づけです。
今回のMR中途採用では、次のような採用方針(外生原理)を設定しました。
この外生原理を5軸のルーブリックRへ落とし込んだ上で、各軸のNorm N(生成評価値)を導出する際、重要な設計判断がありました。「軸ごとの重み(配点比率)」と「軸ごとの絶対的な到達目標(0〜100点の水準)」を、明確に分けて定義することです。
| 評価軸 | 配点(軸間の重み) | Norm N(生成評価値) |
|---|---|---|
| 専門性・疾患知識 | 20% | 70点 |
| 渉外・関係構築力 | 25% | 85点 |
| コンプライアンス意識 | 25% | 85点 |
| 成果創出力 | 15% | 65点 |
| 組織適応・自律性 | 15% | 70点 |
「重視する軸ほど配点が高い」だけでなく、「重視する軸ほど、ルーブリックRから生成される評価値そのものも高くなる」という二重の反映です。この生成評価値Nは合否を分ける足切り線ではなく、SEGTのV*(収束後スコア)がNと評価者集団Mのスコアを連続的にすり合わせて算出するための基準点です。
② ルーブリック定義——評価者が共通言語で採点できるように
Norm Nを生成する土台となるルーブリックRそのものが曖昧では、評価者ごとに解釈がバラバラになってしまいます。5段階のアンカー基準を作成し、各軸で「どのような記述があれば何点相当か」を具体化しました。例えば「渉外・関係構築力」は次の通りです。
- 5(卓越):KOLとの継続的信頼関係構築の具体エピソードと成果が明確
- 4(優良):医療機関担当者との関係構築実績が具体的に記述されている
- 3(標準):関係構築への言及はあるが成果や継続性が不明確
- 2(要改善):対人折衝経験はあるが医療機関特有の文脈がない
- 1(不足):渉外・関係構築の記述がほぼない
③ 審査員構成(評価者集団M)——4つのペルソナから2つのパネルを編成する
SEGTのもう一つの構成要素が、評価者集団Mです。評価者を「複数の視点を持つペルソナ」として設計できます。今回は次の4ペルソナを定義しました。
| ペルソナ | 重視する視点 |
|---|---|
| 採用マネージャー型 | KPI・即戦力 |
| コンプライアンス責任者型 | 規制遵守 |
| 現場MRリーダー型 | 渉外・KOL対応の実務適性 |
| 育成観点型 | 伸びしろ・学習意欲 |
この4名を、目的の異なる2つの3名パネル(M)に再編成しました。コンプライアンス責任者型と現場MRリーダー型はMR職の核となる評価軸のため両パネルに共通させ、採用マネージャー型と育成観点型だけを入れ替えています。
組1|即戦力重視パネル(M₁):採用マネージャー型+コンプライアンス責任者型+現場MRリーダー型
組2|将来性重視パネル(M₂):コンプライアンス責任者型+現場MRリーダー型+育成観点型
ここまでで、外生原理から生成したNorm N(①)と、評価者集団M(③)が揃いました。SEGTはこの(M, N)を収束アルゴリズムΦに入力し、標準評価値V*を算出します。Nは共通のまま、Mだけを2通りに変えることで、V*がどう変化するかを次に検証します。
④ 即戦力重視vs将来性重視で異なる標準評価結果
MR未経験(他業界の法人営業出身)で、独学によるコンプライアンス理解と学習意欲を強みとする架空の応募者データを、M₁・M₂それぞれのパネルで評価しました。
| 評価軸 | 組1 v*(M₁:即戦力) | 組2 v*(M₂:将来性) | Norm N |
|---|---|---|---|
| 専門性・疾患知識 | 59.17 | 62.73 | 70.00 |
| 渉外・関係構築力 | 73.43 | 77.66 | 85.00 |
| コンプライアンス意識 | 79.88 | 83.12 | 85.00 |
| 成果創出力 | 38.32 | 46.64 | 65.00 |
| 組織適応・自律性 | 61.56 | 65.66 | 70.00 |
| 総合スコア | 62.5 | 67.2 | — |
Nは同一のまま、Mだけを変えたにもかかわらず、総合スコアが62.5点から67.2点へと変化しました。特に成果創出力は実務経験に乏しいという実態を反映してNorm N(65点)を大きく下回っていますが、育成観点型を含むM₂ではその不足がやや緩和されています。
興味深いのは、変化が成果創出力だけに留まらず、5軸すべてでM₂の方が一様に高くなっている点です。「Mを構成する1名の視点が変わると、狙った軸だけでなく評価全体が方向性を持ってシフトする」という挙動は、複数評価者による標準評価の設計において重要な示唆です。
成果創出力では、評価者集団Mの採点平均(Eval avg)がわずか10点前後だったのに対し、外生原理から生成されたNorm Nは65点でした。V*(38.32)はどちらか一方に寄らず、両者から距離を保った中間的な値に収束しています。これは算出の不具合ではなく、「外生原理が求める水準と、Mが見た応募者の実態との間に、埋めがたいギャップが存在する」という事実そのものを、V*が定量的に可視化していると解釈できます。
まとめと今後の展開
- 外生原理をルーブリックRに翻訳し、そこから生成評価値Nを導出するという三段階の設計にすることで、採用方針をより精緻に評価へ反映できる
- 評価者集団Mの構成を変えることで、同じNorm Nに対する標準評価値V*が変化し、その変化は特定軸だけでなく全体的な傾向としても現れうる
- Norm NとEval avgの乖離幅そのものが、応募者の強み・弱みを定量的に語る指標として機能する
今後は、特定軸が著しく水準に届かない場合に総合評価によらず選考の再検討を促す仕組みや、より多様な評価者集団Mでのロバスト性検証など、実運用に向けた設計を継続して検討していきます。
Should Standard Evaluation Change Depending on Who Judges?
SEGT Applied to First-Round MR Hiring Screening
If the panel of evaluators changes, how should the standard evaluation change with it? We tested this question using eval000's SEGT (Standard Evaluation Generation Theory) on a first-round document screening scenario for mid-career Medical Representative (MR) hiring at a pharmaceutical company.
First-round document screening is usually standardized by having multiple evaluators score independently and averaging the results. But different evaluators naturally weigh different things: a hiring manager looking for immediate impact and a development-minded manager looking for potential will read the same résumé differently.
Should that variation simply be averaged away as noise? Or should the evaluation draw on both the company's standard and the genuine differences between evaluator perspectives? We used SEGT to explore this question directly.
① Generating Norm N from the External Principle: "Principle → Rubric R → Generated Value N"
Stated precisely, the hiring policy text itself is not "Norm N." Theoretically, the relationship breaks down into four stages:
A hiring policy such as "hire an MR who can perform immediately" (the external principle) is first translated into rubric R — the evaluation axes and anchor criteria. Norm N is not the principle itself, but rather the generated value derived by reading the external principle through rubric R.
For this case, we set the following hiring policy (external principle).
Translating this principle into a 5-axis rubric R, and then deriving each axis's Norm N, required one key design decision: explicitly separating the weight of each axis (how much it contributes to the total, summing to 100%) from the target level (an absolute 0–100 anchor value).
| Evaluation Axis | Weight | Norm N (generated value) |
|---|---|---|
| Expertise / Disease Knowledge | 20% | 70 |
| Relationship Building | 25% | 85 |
| Compliance Awareness | 25% | 85 |
| Track Record / Results | 15% | 65 |
| Organizational Adaptability | 15% | 70 |
Axes the company prioritizes get both a higher weight and a higher value generated from rubric R — a double emphasis. This generated value N is not a pass/fail cutoff; it is the anchor that SEGT's V* (post-convergence score) continuously reconciles against the evaluator population M's actual scores.
② Rubric Design: A Shared Language for Evaluators
If rubric R — the foundation from which Norm N is generated — is itself vague, evaluators will interpret each axis differently. We built a 5-level anchor rubric so each axis had a concrete definition of what a given score should look like. For "Relationship Building":
- 5 (Excellent): Clear, specific episodes and outcomes of sustained trust-building with KOLs
- 4 (Good): Concrete track record of relationship-building with medical institution staff
- 3 (Standard): Relationship-building is mentioned but outcomes or continuity are unclear
- 2 (Needs Improvement): Has interpersonal experience but lacks healthcare-specific context
- 1 (Insufficient): Almost no mention of relationship-building
③ Evaluator Panel Design (Population M): Two Panels from Four Personas
The other component of SEGT is the evaluator population, M. Evaluators can be designed as personas with distinct perspectives. We defined four:
| Persona | Priority |
|---|---|
| Hiring Manager type | KPIs / immediate impact |
| Compliance Officer type | Regulatory adherence |
| Field MR Leader type | Practical fit for KOL engagement |
| Talent Development type | Growth potential / learning drive |
These four were recombined into two 3-person panels (M). The Compliance Officer and Field MR Leader personas — core to the MR role — appear in both panels; only the Hiring Manager and Talent Development personas are swapped.
Panel 1 | Immediate-Impact Panel (M₁): Hiring Manager + Compliance Officer + Field MR Leader
Panel 2 | Future-Potential Panel (M₂): Compliance Officer + Field MR Leader + Talent Development
At this point we have both Norm N, generated from the external principle (①), and the evaluator population M (③). SEGT feeds this pair (M, N) into a convergence algorithm Φ to compute the standard evaluation value V*. Holding N fixed and varying only M lets us examine how V* shifts.
④ Different Standard Evaluations for Immediate-Impact vs. Future-Potential Panels
We evaluated a fictional applicant — no direct MR experience (corporate sales in another industry), with self-taught compliance knowledge and strong learning motivation as their key strength — under both panels M₁ and M₂.
| Axis | Panel 1 v* (M₁: Immediate) | Panel 2 v* (M₂: Future) | Norm N |
|---|---|---|---|
| Expertise / Disease Knowledge | 59.17 | 62.73 | 70.00 |
| Relationship Building | 73.43 | 77.66 | 85.00 |
| Compliance Awareness | 79.88 | 83.12 | 85.00 |
| Track Record / Results | 38.32 | 46.64 | 65.00 |
| Organizational Adaptability | 61.56 | 65.66 | 70.00 |
| Overall Score | 62.5 | 67.2 | — |
With N held constant and only M changed, the overall score shifted from 62.5 to 67.2. Track Record / Results falls well short of its Norm N (65) under both panels, reflecting the applicant's limited hands-on experience — but the shortfall is somewhat softened under M₂, which includes the Talent Development persona.
Notably, the shift wasn't confined to Track Record — all five axes scored uniformly higher under M₂. This suggests that swapping a single evaluator's perspective within M can shift the evaluation's overall tenor, not just the axis it was intended to affect — an important consideration when designing multi-evaluator standard evaluation systems.
For Track Record / Results, the evaluator population M's raw average (Eval avg) was only around 10, against a Norm N of 65 generated from the external principle. Rather than settling near either value, V* (38.32) converged to a point clearly distant from both. This isn't a computational glitch — it's V* quantitatively surfacing a genuine, hard-to-close gap between what the external principle requires and what M observed in the applicant's actual profile.
Conclusion and Next Steps
- Translating the external principle into rubric R, then deriving the generated value N from it, allows a hiring policy to be reflected in evaluation with much greater precision
- Changing the composition of evaluator population M shifts the standard evaluation value V* for the same Norm N, and that shift can appear as a broad tendency rather than being confined to a single axis
- The gap between Norm N and Eval avg itself functions as a quantitative signal of an applicant's strengths and weaknesses
Next steps include designing mechanisms that flag candidates for reconsideration when a specific axis falls far short of target regardless of overall score, and testing robustness across a wider range of evaluator population M configurations.
