[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: [Xen-devel] [PATCH v5 3/8] qspinlock, x86: Add x86 specific optimization for 2 contending tasks

To: Peter Zijlstra <peterz@xxxxxxxxxxxxx>
From: Waiman Long <waiman.long@xxxxxx>
Date: Thu, 27 Feb 2014 15:42:19 -0500
Cc: Jeremy Fitzhardinge <jeremy@xxxxxxxx>, Raghavendra K T <raghavendra.kt@xxxxxxxxxxxxxxxxxx>, Boris Ostrovsky <boris.ostrovsky@xxxxxxxxxx>, virtualization@xxxxxxxxxxxxxxxxxxxxxxxxxx, Andi Kleen <andi@xxxxxxxxxxxxxx>, "H. Peter Anvin" <hpa@xxxxxxxxx>, Michel Lespinasse <walken@xxxxxxxxxx>, Alok Kataria <akataria@xxxxxxxxxx>, linux-arch@xxxxxxxxxxxxxxx, x86@xxxxxxxxxx, Ingo Molnar <mingo@xxxxxxxxxx>, Scott J Norton <scott.norton@xxxxxx>, xen-devel@xxxxxxxxxxxxxxxxxxxx, "Paul E. McKenney" <paulmck@xxxxxxxxxxxxxxxxxx>, Alexander Fyodorov <halcy@xxxxxxxxx>, Arnd Bergmann <arnd@xxxxxxxx>, Daniel J Blueman <daniel@xxxxxxxxxxxxx>, Rusty Russell <rusty@xxxxxxxxxxxxxxx>, Oleg Nesterov <oleg@xxxxxxxxxx>, Steven Rostedt <rostedt@xxxxxxxxxxx>, Chris Wright <chrisw@xxxxxxxxxxxx>, George Spelvin <linux@xxxxxxxxxxx>, Thomas Gleixner <tglx@xxxxxxxxxxxxx>, Aswin Chandramouleeswaran <aswin@xxxxxx>, Chegu Vinod <chegu_vinod@xxxxxx>, Linus Torvalds <torvalds@xxxxxxxxxxxxxxxxxxxx>, linux-kernel@xxxxxxxxxxxxxxx, David Vrabel <david.vrabel@xxxxxxxxxx>, Paolo Bonzini <pbonzini@xxxxxxxxxx>, Andrew Morton <akpm@xxxxxxxxxxxxxxxxxxxx>, Tim Chen <tim.c.chen@xxxxxxxxxxxxxxx>
Delivery-date: Thu, 27 Feb 2014 20:42:49 +0000
List-id: Xen developer discussion <xen-devel.lists.xen.org>

On 02/26/2014 11:20 AM, Peter Zijlstra wrote:

You don't happen to have a proper state diagram for this thing do you?

I suppose I'm going to have to make one; this is all getting a bit
unwieldy, and those xchg() + fixup things are hard to read.

I don't have a state diagram on hand, but I will add more comments todescribe the 4 possible cases and how to handle them.


On Wed, Feb 26, 2014 at 10:14:23AM -0500, Waiman Long wrote:

+static inline int queue_spin_trylock_quick(struct qspinlock *lock, int qsval)
+{
+       union arch_qspinlock *qlock = (union arch_qspinlock *)lock;
+       u16                  old;
+
+       /*
+        * Fall into the quick spinning code path only if no one is waiting
+        * or the lock is available.
+        */
+       if (unlikely((qsval != _QSPINLOCK_LOCKED)&&
+                    (qsval != _QSPINLOCK_WAITING)))
+               return 0;
+
+       old = xchg(&qlock->lock_wait, _QSPINLOCK_WAITING|_QSPINLOCK_LOCKED);
+
+       if (old == 0) {
+               /*
+                * Got the lock, can clear the waiting bit now
+                */
+               smp_u8_store_release(&qlock->wait, 0);


So we just did an atomic op, and now you're trying to optimize this
write. Why do you need a whole byte for that?

Surely a cmpxchg loop with the right atomic op can't be _that_ much
slower? Its far more readable and likely avoids that steal fail below as
well.

At low contention level, atomic operations that requires a lock prefixare the major contributor to the total execution times. I saw estimateonline that the time to execute a lock prefix instruction can easily be50X longer than a regular instruction that can be pipelined. That is whyI try to do it with as few lock prefix instructions as possible. If Ihave to do an atomic cmpxchg, it probably won't be faster than theregular qspinlock slowpath.

Given that speed at low contention level which is the common case isimportant to get this patch accepted, I have to do what I can to make itrun as far as possible for this 2 contending task case.

+               return 1;
+       } else if (old == _QSPINLOCK_LOCKED) {
+try_again:
+               /*
+                * Wait until the lock byte is cleared to get the lock
+                */
+               do {
+                       cpu_relax();
+               } while (ACCESS_ONCE(qlock->lock));
+               /*
+                * Set the lock bit&  clear the waiting bit
+                */
+               if (cmpxchg(&qlock->lock_wait, _QSPINLOCK_WAITING,
+                          _QSPINLOCK_LOCKED) == _QSPINLOCK_WAITING)
+                       return 1;
+               /*
+                * Someone has steal the lock, so wait again
+                */
+               goto try_again;

That's just a fail.. steals should not ever be allowed. It's a fair lock
after all.

The code is unfair, but this unfairness help it to run faster thanticket spinlock in this particular case. And the regular qspinlockslowpath is fair. A little bit of unfairness in this particular casehelps its speed.

+       } else if (old == _QSPINLOCK_WAITING) {
+               /*
+                * Another task is already waiting while it steals the lock.
+                * A bit of unfairness here won't change the big picture.
+                * So just take the lock and return.
+                */
+               return 1;
+       }
+       /*
+        * Nothing need to be done if the old value is
+        * (_QSPINLOCK_WAITING | _QSPINLOCK_LOCKED).
+        */
+       return 0;
+}

@@ -296,6 +478,9 @@ void queue_spin_lock_slowpath(struct qspinlock *lock, int 
qsval)
                return;
        }

+#ifdef queue_code_xchg
+       prev_qcode = queue_code_xchg(lock, my_qcode);
+#else
        /*
         * Exchange current copy of the queue node code
         */
@@ -329,6 +514,7 @@ void queue_spin_lock_slowpath(struct qspinlock *lock, int 
qsval)
        } else
                prev_qcode&= ~_QSPINLOCK_LOCKED;    /* Clear the lock bit */
        my_qcode&= ~_QSPINLOCK_LOCKED;
+#endif /* queue_code_xchg */

        if (prev_qcode) {
                /*

That's just horrible.. please just make the entire #else branch another
version of that same queue_code_xchg() function.


OK, I will wrap it in another function.

Regards,
Longman

_______________________________________________
Xen-devel mailing list
Xen-devel@xxxxxxxxxxxxx
http://lists.xen.org/xen-devel

Follow-Ups:
- Re: [Xen-devel] [PATCH v5 3/8] qspinlock, x86: Add x86 specific optimization for 2 contending tasks
  - From: Peter Zijlstra

References:
- [Xen-devel] [PATCH v5 3/8] qspinlock, x86: Add x86 specific optimization for 2 contending tasks
  - From: Waiman Long
- Re: [Xen-devel] [PATCH v5 3/8] qspinlock, x86: Add x86 specific optimization for 2 contending tasks
  - From: Peter Zijlstra

Prev by Date: Re: [Xen-devel] [PATCH v5 1/8] qspinlock: Introducing a 4-byte queue spinlock implementation
Next by Date: Re: [Xen-devel] [PATCH RFC v5 7/8] pvqspinlock, x86: Add qspinlock para-virtualization support
Previous by thread: Re: [Xen-devel] [PATCH v5 3/8] qspinlock, x86: Add x86 specific optimization for 2 contending tasks
Next by thread: Re: [Xen-devel] [PATCH v5 3/8] qspinlock, x86: Add x86 specific optimization for 2 contending tasks
Index(es):
- Date
- Thread

Lists.xenproject.org is hosted with RackSpace, monitoring our
servers 24x7x365 and backed by RackSpace's Fanatical Support®.