From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1751851AbdJSHbJ (ORCPT ); Thu, 19 Oct 2017 03:31:09 -0400 Received: from mail-eopbgr20135.outbound.protection.outlook.com ([40.107.2.135]:8486 "EHLO EUR02-VE1-obe.outbound.protection.outlook.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S1751711AbdJSHbG (ORCPT ); Thu, 19 Oct 2017 03:31:06 -0400 Authentication-Results: spf=none (sender IP is ) smtp.mailfrom=avagin@virtuozzo.com; Date: Thu, 19 Oct 2017 00:30:52 -0700 From: Andrei Vagin To: Gargi Sharma Cc: linux-kernel@vger.kernel.org, riel@surriel.com, julia.lawall@lip6.fr, akpm@linux-foundation.org, mingo@kernel.org, pasha.tatashin@oracle.com, ktkhai@virtuozzo.com, oleg@redhat.com, ebiederm@xmission.com, hch@infradead.org, lkp@intel.com, tony.luck@intel.com Subject: Re: [v6,1/2] pid: Replace pid bitmap implementation with IDR API Message-ID: <20171019073050.GC29091@outlook.office365.com> References: <1507760379-21662-2-git-send-email-gs051095@gmail.com> MIME-Version: 1.0 Content-Type: text/plain; charset=koi8-r Content-Disposition: inline In-Reply-To: <1507760379-21662-2-git-send-email-gs051095@gmail.com> User-Agent: Mutt/1.8.3 (2017-05-23) X-Originating-IP: [73.140.212.29] X-ClientProxiedBy: LO2P265CA0070.GBRP265.PROD.OUTLOOK.COM (2603:10a6:600:60::34) To HE1PR08MB0746.eurprd08.prod.outlook.com (2a01:111:e400:59b1::12) X-MS-PublicTrafficType: Email X-MS-Office365-Filtering-Correlation-Id: f6977ae1-bbc2-429a-74d7-08d516c35997 X-Microsoft-Antispam: UriScan:;BCL:0;PCL:0;RULEID:(22001)(2017030254152)(2017052603199)(201703131423075)(201703031133081)(201702281549075);SRVR:HE1PR08MB0746; X-Microsoft-Exchange-Diagnostics: 1;HE1PR08MB0746;3:lW2YKxXhlN/2v5BBa0bzeb59atNFFxeXGHC7AlwCiCvB6LxxFLWYXgDf4Mc1D5FC7WywzCnj/9oAl8DnMyKfWq1ZHLgne6o0QE8fQ1Ef99SYQ160fLhXdTPtvMmHaJ1cVYNQ/1ep+LCJ8cC8vOK/QCt1UPr1xOYW7YBx+FPalb2VMcYyYVPg/tXMyE4JYLcxJFhB0NsWXr4dufEqyrZ/qRU0d6bCHaiKMlT/lljxUlD20ew+9LHyh1xJPvcEoi1N;25:Srk5PZbQs1xH4fLobr58KYRByouJat8lalUg5dForITOo7U+qjmkyAlb9nZoZLvDBpG82PUDABRWNR04tPE6QtffWzfUDXNIeFjdtnZILAeGoEnJ4vlGJL47ybK9kLX8gtu03B6bUM4camLAMrCiMp+01uwNbBAQaFuaQwdHCimuUm/SPR0jgD8U8+onu42JUddzD443PX+GKxGTWZaNnQFmJnmkUkeaASLodwOqMCCIZlhGeSwdjXxC926bmsnShI+/QAaMZnpXV8yPvYVkSgvTxW/Ff88V+f7eYdWU59jd7l93VnRbu7z1rddu3Qk7ZvCLeq+u7lL2BYXdXZbU1g==;31:xFVlyEnJQ9FuilF18ulUnfWdLbOxREKowZlTCEqyIrWSc0k9nJRIzVHcwMbvQdSzSuRwvGWR9R7iEipRWfpUCLoDGdo7xYodd7KuTY2ATjHMakQM3wLvXl2GZvxBoIKSJweHsK+cjT6oZE8866XyU5YJwa8FqhY0v0/pTWTir1KAo/Xof8EsitbcQNEvxK6uyslAX1w+q1xEUnNMSzw9QiYd7TjhWoI6yCNK57gaRqE= X-MS-TrafficTypeDiagnostic: HE1PR08MB0746: X-Microsoft-Exchange-Diagnostics: 1;HE1PR08MB0746;20:oXLpiTJv+HEw9kAX4/3TMNfUVxG9ncoXvvXfh0P+Ch/ZwnVvAIYTSVNhtaXey5TBzQXCoAqZ1gMIWLOwEtpqGAj93XrEgKSRFPOuZrqrrudCzJQhYmbKTJlU7S11k076ddBJQvq1Fhpcd6ISJBT1keUoiTeGeT3AEmB6Iz9TDO5HRkxY5XNB61PZMw5kOlviza5mkxZ4wMM3i0Xhgw3fj+ZHu9UbO2zdYnb7NL6ojq6JS0p0eh746t0gFqajVmkRy2iCx+uctjNU4IVYt3KbftfeHq+wTpkzbbZ2bMOMNbvM02biEwG2vsA8GsWESoQ9tS4423AVzn2fr5GitXQryDxvevZbwuojytgmaLbQelbf4AktQsZImecMNBAusjDkfEYAGn0RsHpOVU+W6TPBS7t5L2zv96r/iB1ONWx2OiU=;4:s87m/4wBc1YqfWHZIl+jubhJoxF6IwEFEHYOgVmvHuI//DmH6IauvbgKaM8rYvjNBHhJh7nMJUQMVoTXcpyBLn5Md1AkgEf7gkWKU6rAJPudRfgInTSr4/HLDyAO6HfsX0WUdyfqCicys0eAlSTHnJ8BZeO2szsA3nhrjjcNFEDSAQv3/36gRctJl+YTRBSDQHzrwCNMl7jLdbf1a7pTuNoG1epD72CayFqmEAhvWMxKqx/G4fjDVoHnFXccii8GvaJtVDW0dKcNxE0jUBsWntmW40Ni6QQaAHemqK/vhDw= X-Exchange-Antispam-Report-Test: UriScan:(788757137089); X-Microsoft-Antispam-PRVS: X-Exchange-Antispam-Report-CFA-Test: BCL:0;PCL:0;RULEID:(100000700101)(100105000095)(100000701101)(100105300095)(100000702101)(100105100095)(6040450)(2401047)(8121501046)(5005006)(100000703101)(100105400095)(10201501046)(93006095)(93001095)(3002001)(6041248)(201703131423075)(201702281528075)(201703061421075)(201703061406153)(20161123560025)(20161123564025)(20161123562025)(20161123558100)(20161123555025)(6072148)(201708071742011)(100000704101)(100105200095)(100000705101)(100105500095);SRVR:HE1PR08MB0746;BCL:0;PCL:0;RULEID:(100000800101)(100110000095)(100000801101)(100110300095)(100000802101)(100110100095)(100000803101)(100110400095)(100000804101)(100110200095)(100000805101)(100110500095);SRVR:HE1PR08MB0746; X-Forefront-PRVS: 0465429B7F X-Forefront-Antispam-Report: SFV:NSPM;SFS:(10019020)(6009001)(346002)(376002)(199003)(24454002)(189002)(5660300001)(83506001)(2950100002)(6916009)(16526018)(105586002)(8936002)(50986999)(81156014)(81166006)(2906002)(97736004)(69596002)(7416002)(6116002)(3846002)(58126008)(68736007)(23686003)(16586007)(106356001)(1076002)(7736002)(316002)(4326008)(53936002)(39060400002)(9686003)(6666003)(305945005)(50466002)(55016002)(66066001)(33656002)(6246003)(478600001)(101416001)(47776003)(1411001)(575784001)(86362001)(189998001)(8676002)(25786009)(53416004)(6506006)(54356999)(76176999)(229853002)(18370500001);DIR:OUT;SFP:1102;SCL:1;SRVR:HE1PR08MB0746;H:outlook.office365.com;FPR:;SPF:None;PTR:InfoNoRecords;A:1;MX:1;LANG:en; X-Microsoft-Exchange-Diagnostics: =?koi8-r?Q?1;HE1PR08MB0746;23:zn0CLy7EghgpGSiTwau9WWMqZTe5mK35jI6eSiUTNQ9?= =?koi8-r?Q?hF3e7fugTae4oq8Fb0PlcCGuBBns8fSZ2BLKx4bqV6cu8p+nfHs4IFxyP9UpEL?= =?koi8-r?Q?6tpMLtMWWTBslUbn6Nwj/4dYgyWbCBkApvLwl8gFZ/awUry8lir/SOhu+HjUhU?= =?koi8-r?Q?qwEDBYA+3q/E2zTjVB8LCUVhhUis+3VgqofiUgXbvwlkETbNJ2Ci9RDfT+ZcVt?= =?koi8-r?Q?a3bdPQVZumrABpUfvJqjcS6OU+5VBHnmF/pAtHsSoiFqUskqrdw/+YkcHrJbDk?= =?koi8-r?Q?0JuIzL9D+NOgoGOTI3pUjudnZqAM5vhnZn1/oSyru9cO0fM2r7MA/k8l0J7ujI?= =?koi8-r?Q?ll83PDhxSoXapXyaESKSwthl0GBfM0nnemd3nWMeJuhOxTOOxuHMO4RV38oq+L?= =?koi8-r?Q?JEXDC56HRW+4QsfUIhMyGwvVsaE1yvCgnHMUrSSeYA344nIZltqIiAS/CcLe3b?= =?koi8-r?Q?8k0PvnfrXyLFAFr1Z+Q6/T5hJ/io9+tvAbe6jbhTyIRrVonRclONrYib1Ne+dW?= =?koi8-r?Q?ITDm5dg3hfRRPdjfUjCv4pUrQpoIx2UJALaqXwlVTubbWC04WrmHWocT7WIjHT?= =?koi8-r?Q?XZUt4n2e6UyLi8uUsUzZ9qpGRsjlksY+O42riMIPHPfQoMLdzcOyM50370vVzs?= =?koi8-r?Q?AtR0yprAp5WxO3bmTKOHUFCi7ia1vqmJwmA4nlSG4MaIhrec0MbjzxiErQOfff?= =?koi8-r?Q?uiPg65CKwYKXL4Mqs+n3LDjI6LsXB6QvIZ9AcNi9sbdW28TPa9W+fUwZpMqZAI?= =?koi8-r?Q?UqlfLWCTGLUoWYaM47xEYYmLeN7DuvvikzYKqcZ+KrQ4oC8rvT1ymKj5x6cCY8?= =?koi8-r?Q?PGZWKr50VEYM4zFCbZgpJwqBiaPfOgwu6PE3Fu2gwYAK/s31hNWr/WeBBCC0iJ?= =?koi8-r?Q?IS+x4ao/rg3pwQy7eMp8QdggWl2HhhdnMS9UO9obihb2ZsV6SIc/x1gSYROuOn?= =?koi8-r?Q?xas1gSgGX2m4r+LN1CYSjlaojuJrFylMIvEtATRn9o1B462BJnAeMF8Lv77pAp?= =?koi8-r?Q?w/pL4hJORv8OvehxkBcSmBzYcNkv1BzboFGiMnWQbrezSE/173hj5XBzVZ0WIQ?= =?koi8-r?Q?RBVVvOiYYvofVub1dr4B1txl8enHnHQ8gr6sxQ7BOK5YDnUeMklXsKDLKPnOWg?= =?koi8-r?Q?F6gDLgh3DWbY0IZGzvecDkkt19K8/LS5QV6ExlYf+W352uVJ5ptkVXteJ5mIwB?= =?koi8-r?Q?LiEVVtRapPaTKqGf4HLY9tXL9L1K/DMgkD5qZ5ac8r39CnFMBq5tsaLPsFKHrh?= =?koi8-r?Q?yN55aciJ756zh9EWhow=3D=3D?= X-Microsoft-Exchange-Diagnostics: 1;HE1PR08MB0746;6:PZrcQP1GDiSo7mYSMUAigPqlGyIYUsAEeqTB5ozo8B3HrXI19QvdDs6CnclPY5YxFG3ZMrTCh5D9LzR0pi6c98ZNlN72+p8yaHDM+hlPGQtNvKCDLeevWJoT+hgaXqOf71LLsiMPGJ7LG7Qup9BhoLXI4k6sEGu3nf0aX3mAjxjdlNzP+aLVkkbkG4+96m8ZOFzgZb8+jr4wE2lqRBKY128+Z6Ot8BeDxJmqpWu90uOIVEYpJNRN44q6OZkq3hgwjaelmV2wIxIeb8VlpGYwFMifFVWcpjmZNd5rsCGt5h5te2qwbMiLO5pNFRVdSgNSVTuUHaUEcXZqtOYSGQhTmw==;5:sgIj8qpvmhsu+mdu9HXseYpfxmez00xcWJi4BFXsTy0RSkDCXrxzELsvuYSmAd8TwfX07qPe3azZwUyMtYGaUP+avt9Z/+l8XKNDvO9AQtunKteBkBUGhNqB6S5MWrBRuWKzzznT0GUMBYHrd38EzA==;24:AP1SW8o79s8PrO+DktjG/6gspKK7i7wcKaDEKtXs7eBHcj9tIZ8QM5LcTtb7eIidlAdToKLq0sJb58L23Ksc7yB0XuozqUjgMw2Pqmx66KA=;7:wJRbzKN3/JtiiqqjlO35+4dOkGW7IUvOZoNWg2CTDwAs4j5/WDbQGvKe9M9ySOuxeR/opTlGhKuSuaNb9eVOKnP57e26hApgFzRZ4oaTaNGksMQPaQjsNYdPuUGaUfRHUTh9H1ycsllXjj2+WsY+7QnSaGp4+BMkeJjYV4JcOM6wWyTlthUBq6XxKf7CHY6+yfOwljXfdXspAF+RDp9Sgdbig/+nmSkKWGPA83nhMoM= SpamDiagnosticOutput: 1:99 SpamDiagnosticMetadata: NSPM X-Microsoft-Exchange-Diagnostics: 1;HE1PR08MB0746;20:NGtLQ3vzrc7RyxtFi/Vv6cbjWDiLBOdFCITI4GuotVn7IJsYNAHbUO0semfLknZGfuO5zcigVPtgoqyhrHSRlcHGLOxxCE/+fSbz9H006govGnV3cSeGlkdrBOlm3CqVXhiBmyFP3Vi1bgW7ZAsQ8m9AXwoEIXYizCyEC8Rz9Os= X-OriginatorOrg: virtuozzo.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 19 Oct 2017 07:30:58.7042 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 0bc7f26d-0264-416e-a6fc-8352af79c58f X-MS-Exchange-Transport-CrossTenantHeadersStamped: HE1PR08MB0746 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi Gargi, This patch breaks CRIU, because it changes a meaning of ns_last_pid. ========================== Run zdtm/static/env00 in h ========================== DEP env00.d CC env00.o LINK env00 Start test ./env00 --pidfile=env00.pid --outfile=env00.out --envname=ENV_00_TEST Run criu dump Run criu restore =[log]=> dump/zdtm/static/env00/52/1/restore.log ------------------------ grep Error ------------------------ (00.000587) No mountpoints-6.img image (00.000593) mnt: Reading mountpoint images (id 6 pid 52) (00.000653) Forking task with 52 pid (flags 0x0) (00.007568) PID: real 51 virt 52 (00.010363) 52: Error (criu/cr-restore.c:1787): Pid 51 do not match expected 52 (task 52) (00.010474) Error (criu/cr-restore.c:2449): Restoring FAILED. ------------------------ ERROR OVER ------------------------ Before this patch, ns_last_pid contains a pid of a last process. With this patch, it contains a pid of a "next" process. In CRIU we use ns_last_pid to restore a process with a specified pid, and now this logic is broken: $ uname -a Linux laptop 4.11.11-200.fc25.x86_64 #1 SMP Mon Jul 17 17:41:12 UTC 2017 x86_64 x86_64 x86_64 GNU/Linux $ echo 19999 > /proc/sys/kernel/ns_last_pid && sh -c 'echo $$' 20000 $ uname -a Linux fc24 4.14.0-rc5-next-20171018 #1 SMP Wed Oct 18 23:52:43 PDT 2017 x86_64 x86_64 x86_64 GNU/Linux $ echo 19999 > /proc/sys/kernel/ns_last_pid && sh -c 'echo $$' 19999 Thanks, Andrei On Wed, Oct 11, 2017 at 06:19:38PM -0400, Gargi Sharma wrote: > This patch replaces the current bitmap implemetation for > Process ID allocation. Functions that are no longer required, > for example, free_pidmap(), alloc_pidmap(), etc. are removed. > The rest of the functions are modified to use the IDR API. > The change was made to make the PID allocation less complex by > replacing custom code with calls to generic API. > > Signed-off-by: Gargi Sharma > Reviewed-by: Rik van Riel > --- > arch/powerpc/platforms/cell/spufs/sched.c | 2 +- > fs/proc/loadavg.c | 2 +- > include/linux/pid_namespace.h | 14 +-- > init/main.c | 2 +- > kernel/pid.c | 201 ++++++------------------------ > kernel/pid_namespace.c | 44 +++---- > 6 files changed, 57 insertions(+), 208 deletions(-) > > diff --git a/arch/powerpc/platforms/cell/spufs/sched.c b/arch/powerpc/platforms/cell/spufs/sched.c > index 1fbb5da..e47761c 100644 > --- a/arch/powerpc/platforms/cell/spufs/sched.c > +++ b/arch/powerpc/platforms/cell/spufs/sched.c > @@ -1093,7 +1093,7 @@ static int show_spu_loadavg(struct seq_file *s, void *private) > LOAD_INT(c), LOAD_FRAC(c), > count_active_contexts(), > atomic_read(&nr_spu_contexts), > - task_active_pid_ns(current)->last_pid); > + idr_get_cursor(&task_active_pid_ns(current)->idr)); > return 0; > } > > diff --git a/fs/proc/loadavg.c b/fs/proc/loadavg.c > index 983fce5..ba3d0e2 100644 > --- a/fs/proc/loadavg.c > +++ b/fs/proc/loadavg.c > @@ -23,7 +23,7 @@ static int loadavg_proc_show(struct seq_file *m, void *v) > LOAD_INT(avnrun[1]), LOAD_FRAC(avnrun[1]), > LOAD_INT(avnrun[2]), LOAD_FRAC(avnrun[2]), > nr_running(), nr_threads, > - task_active_pid_ns(current)->last_pid); > + idr_get_cursor(&task_active_pid_ns(current)->idr)); > return 0; > } > > diff --git a/include/linux/pid_namespace.h b/include/linux/pid_namespace.h > index b09136f..f4db4a7 100644 > --- a/include/linux/pid_namespace.h > +++ b/include/linux/pid_namespace.h > @@ -9,15 +9,8 @@ > #include > #include > #include > +#include > > -struct pidmap { > - atomic_t nr_free; > - void *page; > -}; > - > -#define BITS_PER_PAGE (PAGE_SIZE * 8) > -#define BITS_PER_PAGE_MASK (BITS_PER_PAGE-1) > -#define PIDMAP_ENTRIES ((PID_MAX_LIMIT+BITS_PER_PAGE-1)/BITS_PER_PAGE) > > struct fs_pin; > > @@ -29,9 +22,8 @@ enum { /* definitions for pid_namespace's hide_pid field */ > > struct pid_namespace { > struct kref kref; > - struct pidmap pidmap[PIDMAP_ENTRIES]; > + struct idr idr; > struct rcu_head rcu; > - int last_pid; > unsigned int nr_hashed; > struct task_struct *child_reaper; > struct kmem_cache *pid_cachep; > @@ -105,6 +97,6 @@ static inline int reboot_pid_ns(struct pid_namespace *pid_ns, int cmd) > > extern struct pid_namespace *task_active_pid_ns(struct task_struct *tsk); > void pidhash_init(void); > -void pidmap_init(void); > +void pid_idr_init(void); > > #endif /* _LINUX_PID_NS_H */ > diff --git a/init/main.c b/init/main.c > index 0ee9c686..9f4db20 100644 > --- a/init/main.c > +++ b/init/main.c > @@ -667,7 +667,7 @@ asmlinkage __visible void __init start_kernel(void) > if (late_time_init) > late_time_init(); > calibrate_delay(); > - pidmap_init(); > + pid_idr_init(); > anon_vma_init(); > acpi_early_init(); > #ifdef CONFIG_X86 > diff --git a/kernel/pid.c b/kernel/pid.c > index 020dedb..0ce5936 100644 > --- a/kernel/pid.c > +++ b/kernel/pid.c > @@ -39,6 +39,7 @@ > #include > #include > #include > +#include > > #define pid_hashfn(nr, ns) \ > hash_long((unsigned long)nr + (unsigned long)ns, pidhash_shift) > @@ -53,14 +54,6 @@ int pid_max = PID_MAX_DEFAULT; > int pid_max_min = RESERVED_PIDS + 1; > int pid_max_max = PID_MAX_LIMIT; > > -static inline int mk_pid(struct pid_namespace *pid_ns, > - struct pidmap *map, int off) > -{ > - return (map - pid_ns->pidmap)*BITS_PER_PAGE + off; > -} > - > -#define find_next_offset(map, off) \ > - find_next_zero_bit((map)->page, BITS_PER_PAGE, off) > > /* > * PID-map pages start out as NULL, they get allocated upon > @@ -70,10 +63,7 @@ static inline int mk_pid(struct pid_namespace *pid_ns, > */ > struct pid_namespace init_pid_ns = { > .kref = KREF_INIT(2), > - .pidmap = { > - [ 0 ... PIDMAP_ENTRIES-1] = { ATOMIC_INIT(BITS_PER_PAGE), NULL } > - }, > - .last_pid = 0, > + .idr = IDR_INIT, > .nr_hashed = PIDNS_HASH_ADDING, > .level = 0, > .child_reaper = &init_task, > @@ -101,138 +91,6 @@ EXPORT_SYMBOL_GPL(init_pid_ns); > > static __cacheline_aligned_in_smp DEFINE_SPINLOCK(pidmap_lock); > > -static void free_pidmap(struct upid *upid) > -{ > - int nr = upid->nr; > - struct pidmap *map = upid->ns->pidmap + nr / BITS_PER_PAGE; > - int offset = nr & BITS_PER_PAGE_MASK; > - > - clear_bit(offset, map->page); > - atomic_inc(&map->nr_free); > -} > - > -/* > - * If we started walking pids at 'base', is 'a' seen before 'b'? > - */ > -static int pid_before(int base, int a, int b) > -{ > - /* > - * This is the same as saying > - * > - * (a - base + MAXUINT) % MAXUINT < (b - base + MAXUINT) % MAXUINT > - * and that mapping orders 'a' and 'b' with respect to 'base'. > - */ > - return (unsigned)(a - base) < (unsigned)(b - base); > -} > - > -/* > - * We might be racing with someone else trying to set pid_ns->last_pid > - * at the pid allocation time (there's also a sysctl for this, but racing > - * with this one is OK, see comment in kernel/pid_namespace.c about it). > - * We want the winner to have the "later" value, because if the > - * "earlier" value prevails, then a pid may get reused immediately. > - * > - * Since pids rollover, it is not sufficient to just pick the bigger > - * value. We have to consider where we started counting from. > - * > - * 'base' is the value of pid_ns->last_pid that we observed when > - * we started looking for a pid. > - * > - * 'pid' is the pid that we eventually found. > - */ > -static void set_last_pid(struct pid_namespace *pid_ns, int base, int pid) > -{ > - int prev; > - int last_write = base; > - do { > - prev = last_write; > - last_write = cmpxchg(&pid_ns->last_pid, prev, pid); > - } while ((prev != last_write) && (pid_before(base, last_write, pid))); > -} > - > -static int alloc_pidmap(struct pid_namespace *pid_ns) > -{ > - int i, offset, max_scan, pid, last = pid_ns->last_pid; > - struct pidmap *map; > - > - pid = last + 1; > - if (pid >= pid_max) > - pid = RESERVED_PIDS; > - offset = pid & BITS_PER_PAGE_MASK; > - map = &pid_ns->pidmap[pid/BITS_PER_PAGE]; > - /* > - * If last_pid points into the middle of the map->page we > - * want to scan this bitmap block twice, the second time > - * we start with offset == 0 (or RESERVED_PIDS). > - */ > - max_scan = DIV_ROUND_UP(pid_max, BITS_PER_PAGE) - !offset; > - for (i = 0; i <= max_scan; ++i) { > - if (unlikely(!map->page)) { > - void *page = kzalloc(PAGE_SIZE, GFP_KERNEL); > - /* > - * Free the page if someone raced with us > - * installing it: > - */ > - spin_lock_irq(&pidmap_lock); > - if (!map->page) { > - map->page = page; > - page = NULL; > - } > - spin_unlock_irq(&pidmap_lock); > - kfree(page); > - if (unlikely(!map->page)) > - return -ENOMEM; > - } > - if (likely(atomic_read(&map->nr_free))) { > - for ( ; ; ) { > - if (!test_and_set_bit(offset, map->page)) { > - atomic_dec(&map->nr_free); > - set_last_pid(pid_ns, last, pid); > - return pid; > - } > - offset = find_next_offset(map, offset); > - if (offset >= BITS_PER_PAGE) > - break; > - pid = mk_pid(pid_ns, map, offset); > - if (pid >= pid_max) > - break; > - } > - } > - if (map < &pid_ns->pidmap[(pid_max-1)/BITS_PER_PAGE]) { > - ++map; > - offset = 0; > - } else { > - map = &pid_ns->pidmap[0]; > - offset = RESERVED_PIDS; > - if (unlikely(last == offset)) > - break; > - } > - pid = mk_pid(pid_ns, map, offset); > - } > - return -EAGAIN; > -} > - > -int next_pidmap(struct pid_namespace *pid_ns, unsigned int last) > -{ > - int offset; > - struct pidmap *map, *end; > - > - if (last >= PID_MAX_LIMIT) > - return -1; > - > - offset = (last + 1) & BITS_PER_PAGE_MASK; > - map = &pid_ns->pidmap[(last + 1)/BITS_PER_PAGE]; > - end = &pid_ns->pidmap[PIDMAP_ENTRIES]; > - for (; map < end; map++, offset = 0) { > - if (unlikely(!map->page)) > - continue; > - offset = find_next_bit((map)->page, BITS_PER_PAGE, offset); > - if (offset < BITS_PER_PAGE) > - return mk_pid(pid_ns, map, offset); > - } > - return -1; > -} > - > void put_pid(struct pid *pid) > { > struct pid_namespace *ns; > @@ -266,7 +124,7 @@ void free_pid(struct pid *pid) > struct upid *upid = pid->numbers + i; > struct pid_namespace *ns = upid->ns; > hlist_del_rcu(&upid->pid_chain); > - switch(--ns->nr_hashed) { > + switch (--ns->nr_hashed) { > case 2: > case 1: > /* When all that is left in the pid namespace > @@ -284,12 +142,11 @@ void free_pid(struct pid *pid) > schedule_work(&ns->proc_work); > break; > } > + > + idr_remove(&ns->idr, upid->nr); > } > spin_unlock_irqrestore(&pidmap_lock, flags); > > - for (i = 0; i <= pid->level; i++) > - free_pidmap(pid->numbers + i); > - > call_rcu(&pid->rcu, delayed_put_pid); > } > > @@ -308,8 +165,29 @@ struct pid *alloc_pid(struct pid_namespace *ns) > > tmp = ns; > pid->level = ns->level; > + > for (i = ns->level; i >= 0; i--) { > - nr = alloc_pidmap(tmp); > + int pid_min = 1; > + > + idr_preload(GFP_KERNEL); > + spin_lock_irq(&pidmap_lock); > + > + /* > + * init really needs pid 1, but after reaching the maximum > + * wrap back to RESERVED_PIDS > + */ > + if (idr_get_cursor(&tmp->idr) > RESERVED_PIDS) > + pid_min = RESERVED_PIDS; > + > + /* > + * Store a null pointer so find_pid_ns does not find > + * a partially initialized PID (see below). > + */ > + nr = idr_alloc_cyclic(&tmp->idr, NULL, pid_min, > + pid_max, GFP_ATOMIC); > + spin_unlock_irq(&pidmap_lock); > + idr_preload_end(); > + > if (nr < 0) { > retval = nr; > goto out_free; > @@ -339,6 +217,8 @@ struct pid *alloc_pid(struct pid_namespace *ns) > for ( ; upid >= pid->numbers; --upid) { > hlist_add_head_rcu(&upid->pid_chain, > &pid_hash[pid_hashfn(upid->nr, upid->ns)]); > + /* Make the PID visible to find_pid_ns. */ > + idr_replace(&upid->ns->idr, pid, upid->nr); > upid->ns->nr_hashed++; > } > spin_unlock_irq(&pidmap_lock); > @@ -350,8 +230,11 @@ struct pid *alloc_pid(struct pid_namespace *ns) > put_pid_ns(ns); > > out_free: > + spin_lock_irq(&pidmap_lock); > while (++i <= ns->level) > - free_pidmap(pid->numbers + i); > + idr_remove(&ns->idr, (pid->numbers + i)->nr); > + > + spin_unlock_irq(&pidmap_lock); > > kmem_cache_free(ns->pid_cachep, pid); > return ERR_PTR(retval); > @@ -553,16 +436,7 @@ EXPORT_SYMBOL_GPL(task_active_pid_ns); > */ > struct pid *find_ge_pid(int nr, struct pid_namespace *ns) > { > - struct pid *pid; > - > - do { > - pid = find_pid_ns(nr, ns); > - if (pid) > - break; > - nr = next_pidmap(ns, nr); > - } while (nr > 0); > - > - return pid; > + return idr_get_next(&ns->idr, &nr); > } > > /* > @@ -578,7 +452,7 @@ void __init pidhash_init(void) > 0, 4096); > } > > -void __init pidmap_init(void) > +void __init pid_idr_init(void) > { > /* Verify no one has done anything silly: */ > BUILD_BUG_ON(PID_MAX_LIMIT >= PIDNS_HASH_ADDING); > @@ -590,10 +464,7 @@ void __init pidmap_init(void) > PIDS_PER_CPU_MIN * num_possible_cpus()); > pr_info("pid_max: default: %u minimum: %u\n", pid_max, pid_max_min); > > - init_pid_ns.pidmap[0].page = kzalloc(PAGE_SIZE, GFP_KERNEL); > - /* Reserve PID 0. We never call free_pidmap(0) */ > - set_bit(0, init_pid_ns.pidmap[0].page); > - atomic_dec(&init_pid_ns.pidmap[0].nr_free); > + idr_init(&init_pid_ns.idr); > > init_pid_ns.pid_cachep = KMEM_CACHE(pid, > SLAB_HWCACHE_ALIGN | SLAB_PANIC | SLAB_ACCOUNT); > diff --git a/kernel/pid_namespace.c b/kernel/pid_namespace.c > index 4918314..5009dbe 100644 > --- a/kernel/pid_namespace.c > +++ b/kernel/pid_namespace.c > @@ -21,6 +21,7 @@ > #include > #include > #include > +#include > > struct pid_cache { > int nr_ids; > @@ -98,7 +99,6 @@ static struct pid_namespace *create_pid_namespace(struct user_namespace *user_ns > struct pid_namespace *ns; > unsigned int level = parent_pid_ns->level + 1; > struct ucounts *ucounts; > - int i; > int err; > > err = -EINVAL; > @@ -117,17 +117,15 @@ static struct pid_namespace *create_pid_namespace(struct user_namespace *user_ns > if (ns == NULL) > goto out_dec; > > - ns->pidmap[0].page = kzalloc(PAGE_SIZE, GFP_KERNEL); > - if (!ns->pidmap[0].page) > - goto out_free; > + idr_init(&ns->idr); > > ns->pid_cachep = create_pid_cachep(level + 1); > if (ns->pid_cachep == NULL) > - goto out_free_map; > + goto out_free_idr; > > err = ns_alloc_inum(&ns->ns); > if (err) > - goto out_free_map; > + goto out_free_idr; > ns->ns.ops = &pidns_operations; > > kref_init(&ns->kref); > @@ -138,17 +136,10 @@ static struct pid_namespace *create_pid_namespace(struct user_namespace *user_ns > ns->nr_hashed = PIDNS_HASH_ADDING; > INIT_WORK(&ns->proc_work, proc_cleanup_work); > > - set_bit(0, ns->pidmap[0].page); > - atomic_set(&ns->pidmap[0].nr_free, BITS_PER_PAGE - 1); > - > - for (i = 1; i < PIDMAP_ENTRIES; i++) > - atomic_set(&ns->pidmap[i].nr_free, BITS_PER_PAGE); > - > return ns; > > -out_free_map: > - kfree(ns->pidmap[0].page); > -out_free: > +out_free_idr: > + idr_destroy(&ns->idr); > kmem_cache_free(pid_ns_cachep, ns); > out_dec: > dec_pid_namespaces(ucounts); > @@ -168,11 +159,9 @@ static void delayed_free_pidns(struct rcu_head *p) > > static void destroy_pid_namespace(struct pid_namespace *ns) > { > - int i; > - > ns_free_inum(&ns->ns); > - for (i = 0; i < PIDMAP_ENTRIES; i++) > - kfree(ns->pidmap[i].page); > + > + idr_destroy(&ns->idr); > call_rcu(&ns->rcu, delayed_free_pidns); > } > > @@ -213,6 +202,7 @@ void zap_pid_ns_processes(struct pid_namespace *pid_ns) > int rc; > struct task_struct *task, *me = current; > int init_pids = thread_group_leader(me) ? 1 : 2; > + struct pid *pid; > > /* Don't allow any more processes into the pid namespace */ > disable_pid_allocation(pid_ns); > @@ -239,20 +229,16 @@ void zap_pid_ns_processes(struct pid_namespace *pid_ns) > * maintain a tasklist for each pid namespace. > * > */ > + rcu_read_lock(); > read_lock(&tasklist_lock); > - nr = next_pidmap(pid_ns, 1); > - while (nr > 0) { > - rcu_read_lock(); > - > - task = pid_task(find_vpid(nr), PIDTYPE_PID); > + nr = 2; > + idr_for_each_entry_continue(&pid_ns->idr, pid, nr) { > + task = pid_task(pid, PIDTYPE_PID); > if (task && !__fatal_signal_pending(task)) > send_sig_info(SIGKILL, SEND_SIG_FORCED, task); > - > - rcu_read_unlock(); > - > - nr = next_pidmap(pid_ns, nr); > } > read_unlock(&tasklist_lock); > + rcu_read_unlock(); > > /* > * Reap the EXIT_ZOMBIE children we had before we ignored SIGCHLD. > @@ -311,7 +297,7 @@ static int pid_ns_ctl_handler(struct ctl_table *table, int write, > * it should synchronize its usage with external means. > */ > > - tmp.data = &pid_ns->last_pid; > + tmp.data = &pid_ns->idr.idr_next; > return proc_dointvec_minmax(&tmp, write, buffer, lenp, ppos); > } >